Team Ai
Datasetpublic

MultilingualUnigramLM/LangMap-TheStack-cpp-100M

LangMap-TheStack-cpp-100M Code finetuning dataset for cpp streamed from bigcode/the-stack. Tokens collected: 100,000,000 (target: 100,000,000) Tokenizer: allenai/OLMo-3-1025-7B Schema: {"text": [...]} (sanitised source code)

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes237downloads
Dataset Card

LangMap-TheStack-cpp-100M

Code finetuning dataset for cpp streamed from bigcode/the-stack.

  • —Tokens collected: 100,000,000 (target: 100,000,000)
  • —Tokenizer: allenai/OLMo-3-1025-7B
  • —Schema: {"text": [...]} (sanitised source code)