Team Ai
Datasetpublic

kernelmachine/open-license-corpus

PubText Welcome to the Open License Corpus (OLC), a 228B token corpus for training permissively-licensed language models. Disclaimer: OLC should not be considered a universally safe-to-use dataset. We encourage users of OLC to consult a legal professional on the suitability of each data source for their application. Dataset Summary Domain Sources Specific License # BPE Tokens (in billions; GPT-NeoX tokenizer) Legal Case Law, Pile of Law (PD subset)… See the full description on the dataset page: https://huggingface.co/datasets/kernelmachine/open-license-corpus.

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
19likes1.6kdownloads
settings

This repository belongs to kernelmachine on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameopen-license-corpus
visibilitypublic
licenceapache-2.0
gatedno
ownerkernelmachine
Account settings
kernelmachine/open-license-corpus · Team Ai