Salesforce/ctrl
1960k
1---2language: en3license: bsd-3-clause4pipeline_tag: text-generation5---6 7# ctrl8 9# Table of Contents10 111. [Model Details](#model-details)122. [Uses](#uses)133. [Bias, Risks, and Limitations](#bias-risks-and-limitations)144. [Training](#training)155. [Evaluation](#evaluation)166. [Environmental Impact](#environmental-impact)177. [Technical Specifications](#technical-specifications)188. [Citation](#citation)199. [Model Card Authors](#model-card-authors)2010. [How To Get Started With the Model](#how-to-get-started-with-the-model)21 22 23# Model Details24 25## Model Description26 27The CTRL model was proposed in [CTRL: A Conditional Transformer Language Model for Controllable Generation](https://arxiv.org/abs/1909.05858) by Nitish Shirish Keskar*, Bryan McCann*, Lav R. Varshney, Caiming Xiong and Richard Socher. It's a causal (unidirectional) transformer pre-trained using language modeling on a very large corpus of ~140 GB of text data with the first token reserved as a control code (such as Links, Books, Wikipedia etc.). The model developers released a model card for CTRL, available [here](https://github.com/salesforce/ctrl/blob/master/ModelCard.pdf).28 29In their [model card](https://github.com/salesforce/ctrl/blob/master/ModelCard.pdf), the developers write: 30 31> The CTRL Language Model analyzed in this card generates text conditioned on control codes that specify domain, style, topics, dates, entities, relationships between entities, plot points, and task-related behavior.32 33- **Developed by:** See [associated paper](https://arxiv.org/abs/1909.05858) from Salesforce Research34- **Model type:** Transformer-based language model35- **Language(s) (NLP):** Primarily English, some German, Spanish, French36- **License:** [BSD 3-Clause](https://github.com/salesforce/ctrl/blob/master/LICENSE.txt); also see [Code of Conduct](https://github.com/salesforce/ctrl)37- **Related Models:** More information needed38 - **Parent Model:** More information needed39- **Resources for more information:** 40 - [Associated paper](https://arxiv.org/abs/1909.05858)41 - [GitHub repo](https://github.com/salesforce/ctrl)42 - [Developer Model Card](https://github.com/salesforce/ctrl/blob/master/ModelCard.pdf)43 - [Blog post](https://blog.salesforceairesearch.com/introducing-a-conditional-transformer-language-model-for-controllable-generation/)44 45# Uses46 47## Direct Use48 49The model is a language model. The model can be used for text generation. 50 51## Downstream Use52 53In their [model card](https://github.com/salesforce/ctrl/blob/master/ModelCard.pdf), the developers write that the primary intended users are general audiences and NLP Researchers, and that the primary intended uses are:54 55> 1. Generating artificial text in collaboration with a human, including but not limited to:56> - Creative writing57> - Automating repetitive writing tasks58> - Formatting specific text types59> - Creating contextualized marketing materials60> 2. Improvement of other NLP applications through fine-tuning (on another task or other data, e.g. fine-tuning CTRL to learn new kinds of language like product descriptions)61> 3. Enhancement in the field of natural language understanding to push towards a better understanding of artificial text generation, including how to detect it and work toward control, understanding, and potentially combating potentially negative consequences of such models.62 63## Out-of-Scope Use64 65In their [model card](https://github.com/salesforce/ctrl/blob/master/ModelCard.pdf), the developers write: 66 67> - CTRL should not be used for generating artificial text without collaboration with a human.68> - It should not be used to make normative or prescriptive claims.69> - This software should not be used to promote or profit from:70> - violence, hate, and division;71> - environmental destruction;72> - abuse of human rights; or73> - the destruction of people's physical and mental health.74 75# Bias, Risks, and Limitations76 77Significant research has explored bias and fairness issues with language models (see, e.g., [Sheng et al. (2021)](https://aclanthology.org/2021.acl-long.330.pdf) and [Bender et al. (2021)](https://dl.acm.org/doi/pdf/10.1145/3442188.3445922)). Predictions generated by the model may include disturbing and harmful stereotypes across protected classes; identity characteristics; and sensitive, social, and occupational groups.78 79In their [model card](https://github.com/salesforce/ctrl/blob/master/ModelCard.pdf), the developers write: 80 81> We recognize the potential for misuse or abuse, including use by bad actors who could manipulate the system to act maliciously and generate text to influence decision-making in political, economic, and social settings. False attribution could also harm individuals, organizations, or other entities. To address these concerns, the model was evaluated internally as well as externally by third parties, including the Partnership on AI, prior to release.82 83> To mitigate potential misuse to the extent possible, we stripped out all detectable training data from undesirable sources. We then redteamed the model and found that negative utterances were often placed in contexts that made them identifiable as such. For example, when using the ‘News’ control code, hate speech could be embedded as part of an apology (e.g. “the politician apologized for saying [insert hateful statement]”), implying that this type of speech was negative. By pre-selecting the available control codes (omitting, for example, Instagram and Twitter from the available domains), we are able to limit the potential for misuse.84 85> In releasing our model, we hope to put it into the hands of researchers and prosocial actors so that they can work to control, understand, and potentially combat the negative consequences of such models. We hope that research into detecting fake news and model-generated content of all kinds will be pushed forward by CTRL. It is our belief that these models should become a common tool so researchers can design methods to guard against malicious use and so the public becomes familiar with their existence and patterns of behavior.86 87See the [associated paper](https://arxiv.org/pdf/1909.05858.pdf) for further discussions about the ethics of LLMs.88 89## Recommendations90 91In their [model card](https://github.com/salesforce/ctrl/blob/master/ModelCard.pdf), the developers write: 92 93> - A recommendation to monitor and detect use will be implemented through the development of a model that will identify CTRLgenerated text.94> - A second recommendation to further screen the input into and output from the model will be implemented through the addition of a check in the CTRL interface to prohibit the insertion into the model of certain negative inputs, which will help control the output that can be generated.95> - The model is trained on a limited number of languages: primarily English and some German, Spanish, French. A recommendation for a future area of research is to train the model on more languages.96 97See the [CTRL-detector GitHub repo](https://github.com/salesforce/ctrl-detector) for more on the detector model.98 99# Training100 101## Training Data102 103In their [model card](https://github.com/salesforce/ctrl/blob/master/ModelCard.pdf), the developers write: 104 105> This model is trained on 140 GB of text drawn from a variety of domains: Wikipedia (English, German, Spanish, and French), Project Gutenberg, submissions from 45 subreddits, OpenWebText, a large collection of news data, Amazon Reviews, Europarl and UN data from WMT (En-De, En-Es, En-Fr), question-answer pairs (no context documents) from ELI5, and the MRQA shared task, which includes Stanford Question Answering Dataset, NewsQA, TriviaQA, SearchQA, HotpotQA, and Natural Questions. See the paper for the full list of training data.106 107## Training Procedure108 109### Preprocessing110 111In the [associated paper](https://arxiv.org/pdf/1909.05858.pdf) the developers write: 112 113> We learn BPE (Sennrich et al., 2015) codes and tokenize the data using fastBPE4, but we use a large vocabulary of roughly 250K tokens. This includes the sub-word tokens necessary to mitigate problems with rare words, but it also reduces the average number of tokens required to generate long text by including most common words. We use English Wikipedia and a 5% split of our collected OpenWebText data for learning BPE codes. We also introduce an unknown token so that during preprocessing we can filter out sequences that contain more than 2 unknown tokens. This, along with the compressed storage for efficient training (TFRecords) (Abadi et al., 2016), reduces our training data to 140 GB from the total 180 GB collected.114 115See the paper for links, references, and further details.116 117### Training118 119In the [associated paper](https://arxiv.org/pdf/1909.05858.pdf) the developers write: 120 121> CTRL has model dimension d = 1280, inner dimension f = 8192, 48 layers, and 16 heads per layer. Dropout with probability 0.1 follows the residual connections in each layer. Token embeddings were tied with the final output embedding layer (Inan et al., 2016; Press & Wolf, 2016).122 123See the paper for links, references, and further details.124 125# Evaluation126 127## Testing Data, Factors & Metrics128 129In their [model card](https://github.com/salesforce/ctrl/blob/master/ModelCard.pdf), the developers write that model performance measures are: 130 131> Performance evaluated on qualitative judgments by humans as to whether the control codes lead to text generated in the desired domain132 133# Environmental Impact134 135Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700). Details are pulled from the [associated paper](https://arxiv.org/pdf/1909.05858.pdf).136 137- **Hardware Type:** TPU v3 Pod138- **Hours used:** Approximately 336 hours (2 weeks)139- **Cloud Provider:** GCP140- **Compute Region:** More information needed141- **Carbon Emitted:** More information needed142 143# Technical Specifications144 145In the [associated paper](https://arxiv.org/pdf/1909.05858.pdf) the developers write: 146 147> CTRL was implemented in TensorFlow (Abadi et al., 2016) and trained with a global batch size of 1024 distributed across 256 cores of a Cloud TPU v3 Pod for 800k iterations. Training took approximately 2 weeks using Adagrad (Duchi et al., 2011) with a linear warmup from 0 to 0.05 over 25k steps. The norm of gradients were clipped to 0.25 as in (Merity et al., 2017). Learning rate decay was not necessary due to the monotonic nature of the Adagrad accumulator. We compared to the Adam optimizer (Kingma & Ba, 2014) while training smaller models, but we noticed comparable convergence rates and significant memory savings with Adagrad. We also experimented with explicit memory-saving optimizers including SM3 (Anil et al., 2019), Adafactor (Shazeer & Stern, 2018), and NovoGrad (Ginsburg et al., 2019) with mixed results.148 149See the paper for links, references, and further details.150 151# Citation152 153**BibTeX:**154 155```bibtex156@article{keskarCTRL2019,157 title={{CTRL - A Conditional Transformer Language Model for Controllable Generation}},158 author={Keskar, Nitish Shirish and McCann, Bryan and Varshney, Lav and Xiong, Caiming and Socher, Richard},159 journal={arXiv preprint arXiv:1909.05858},160 year={2019}161}162```163 164**APA:**165- Keskar, N. S., McCann, B., Varshney, L. R., Xiong, C., & Socher, R. (2019). Ctrl: A conditional transformer language model for controllable generation. arXiv preprint arXiv:1909.05858.166 167# Model Card Authors168 169This model card was written by the team at Hugging Face, referencing the [model card](https://github.com/salesforce/ctrl/blob/master/ModelCard.pdf) released by the developers.170 171# Ethical Considerations172This release is for research purposes only in support of an academic paper. Our models, datasets, and code are not specifically designed or evaluated for all downstream purposes. We strongly recommend users evaluate and address potential concerns related to accuracy, safety, and fairness before deploying this model. We encourage users to consider the common limitations of AI, comply with applicable laws, and leverage best practices when selecting use cases, particularly for high-risk scenarios where errors or misuse could significantly impact people’s lives, rights, or safety. For further guidance on use cases, refer to our AUP and AI AUP.173 174# How to Get Started with the Model175 176Use the code below to get started with the model. See the [Hugging Face ctrl docs](https://huggingface.co/docs/transformers/model_doc/ctrl) for more information.177 178<details>179<summary> Click to expand </summary>180 181```python182>>> from transformers import CTRLTokenizer, CTRLModel183>>> import torch184 185>>> tokenizer = CTRLTokenizer.from_pretrained("ctrl")186>>> model = CTRLModel.from_pretrained("ctrl")187 188>>> # CTRL was trained with control codes as the first token189>>> inputs = tokenizer("Opinion My dog is cute", return_tensors="pt")190>>> assert inputs["input_ids"][0, 0].item() in tokenizer.control_codes.values()191 192>>> outputs = model(**inputs)193 194>>> last_hidden_states = outputs.last_hidden_state195>>> list(last_hidden_states.shape)196```197 198</details>