Team Ai
Datasetpublic

bigcode/commitpack-subset-cf

A subset of CommitPack used for pretraining SantaCoderPack from the OctoPack paper. It focuses on data where the code before + special token + code after fits into 8192 tokens and on 6 languages. The data is in commit format (cf): <commit_before>code_before<commit_message>commit_message<commit_after>code_after.

sourceHugging Faceupdated 3y agoView on Hugging Face
2likes2.4kdownloads
10 commits on main
9f2ee5d3y ago

Update README.md

Muennighoff
bb8c0dc3y ago

Update README.md

Muennighoff
ae63e973y ago

Create README.md

Muennighoff
372ded13y ago

Add

Muennighoff
b3eaf6d3y ago

Add rust

Muennighoff
86b67423y ago

Add

Muennighoff
c090b2c3y ago

Add

Muennighoff
ec47ccf3y ago

Add

Muennighoff
504bcf63y ago

Add python

Muennighoff
72948e73y ago

initial commit

Muennighoff