Team Ai
Apppublic

codeparrot/code-generation-models

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
257likes
codegen.md15 linesDownload Raw Back to datasets
1[Codegen](https://huggingface.co/Salesforce/codegen-16B-mono) is a model for conversational program synthesis, where each problem is interactively solved in multiple steps, each consisting of a natural language specification from the user and a synthesized subprogram from the system. 2 3It was sequentially trained on three datasets:4- [The Pile](https://huggingface.co/datasets/the_pile)5- A 341GB subset of Google’s [BigQuery dataset](https://cloud.google.com/blog/topics/public-datasets/github-on-bigquery-analyze-all-the-open-source-code) of code files from multiple programming languages, keeping only 6: C, C++, Go, Java, JavaScript, and Python 6- 217GB of Python data from GitHub repositories 7 8The second and third datasets used the following preprocessing:9- Exact match deduplication 10- Filtering:11    - Exact match deduplication 12    - Average line length < 100 tokens13    - Maximum line length < 1000 MB14    - Characters being decimal or hexadecimal digits >90% 15