kernelmachine/open-license-corpus
PubText Welcome to the Open License Corpus (OLC), a 228B token corpus for training permissively-licensed language models. Disclaimer: OLC should not be considered a universally safe-to-use dataset. We encourage users of OLC to consult a legal professional on the suitability of each data source for their application. Dataset Summary Domain Sources Specific License # BPE Tokens (in billions; GPT-NeoX tokenizer) Legal Case Law, Pile of Law (PD subset)… See the full description on the dataset page: https://huggingface.co/datasets/kernelmachine/open-license-corpus.
This repository belongs to kernelmachine on Hugging Face.
Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
