kernelmachine/open-license-corpus
PubText Welcome to the Open License Corpus (OLC), a 228B token corpus for training permissively-licensed language models. Disclaimer: OLC should not be considered a universally safe-to-use dataset. We encourage users of OLC to consult a legal professional on the suitability of each data source for their application. Dataset Summary Domain Sources Specific License # BPE Tokens (in billions; GPT-NeoX tokenizer) Legal Case Law, Pile of Law (PD subset)… See the full description on the dataset page: https://huggingface.co/datasets/kernelmachine/open-license-corpus.
Update README.md
update
Update open-license-corpus.py
Update open-license-corpus.py
Update open-license-corpus.py
update
update
update
update
initial commit
