Team Ai
Datasetpublic

kernelmachine/open-license-corpus

PubText Welcome to the Open License Corpus (OLC), a 228B token corpus for training permissively-licensed language models. Disclaimer: OLC should not be considered a universally safe-to-use dataset. We encourage users of OLC to consult a legal professional on the suitability of each data source for their application. Dataset Summary Domain Sources Specific License # BPE Tokens (in billions; GPT-NeoX tokenizer) Legal Case Law, Pile of Law (PD subset)… See the full description on the dataset page: https://huggingface.co/datasets/kernelmachine/open-license-corpus.

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
19likes1.6kdownloads
10 commits on main
384d5e13y ago

Update README.md

kernelmachine
47f7c923y ago

update

kernelmachine
7bcc4873y ago

Update open-license-corpus.py

kernelmachine
1c6e67a3y ago

Update open-license-corpus.py

kernelmachine
943f7283y ago

Update open-license-corpus.py

kernelmachine
0ebf5a93y ago

update

kernelmachine
76fda693y ago

update

kernelmachine
2102ac93y ago

update

kernelmachine
f26634f3y ago

update

kernelmachine
c1e6d6a3y ago

initial commit

kernelmachine