bigcode/starcoder2data-extras
StarCoder2 Extras This is the dataset of extra sources (besides Stack v2 code data) used to train the StarCoder2 family of models. It contains the following subsets: Kaggle (kaggle): Kaggle notebooks from Meta-Kaggle-Code dataset, converted to scripts and prefixed with information on the Kaggle datasets used in the notebook. The file headers have a similar format to Jupyter Structured but the code content is only one single script. StackOverflow (stackoverflow): stackoverflow… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/starcoder2data-extras.
133.7k
1---2dataset_info:3- config_name: arxiv4 features:5 - name: content6 dtype: string7 splits:8 - name: train9 num_bytes: 89223183645.010 num_examples: 155830611 download_size: 4091118687612 dataset_size: 89223183645.013- config_name: documentation14 features:15 - name: project16 dtype: string17 - name: source18 dtype: string19 - name: language20 dtype: string21 - name: content22 dtype: string23 splits:24 - name: train25 num_bytes: 5421472234.026 num_examples: 5973327 download_size: 185345192228 dataset_size: 5421472234.029- config_name: ir_cpp30 features:31 - name: __index_level_0__32 dtype: string33 - name: id34 dtype: string35 - name: content36 dtype: string37 splits:38 - name: train39 num_bytes: 102081135272.040 num_examples: 291665541 download_size: 2604797842242 dataset_size: 102081135272.043- config_name: ir_low_resource44 features:45 - name: __index_level_0__46 dtype: string47 - name: id48 dtype: string49 - name: content50 dtype: string51 - name: size52 dtype: int6453 splits:54 - name: train55 num_bytes: 10383382043.056 num_examples: 39398857 download_size: 246451360358 dataset_size: 10383382043.059- config_name: ir_python60 features:61 - name: id62 dtype: string63 - name: content64 dtype: string65 splits:66 - name: train67 num_bytes: 12446664464.068 num_examples: 15450769 download_size: 303929762570 dataset_size: 12446664464.071- config_name: ir_rust72 features:73 - name: __index_level_0__74 dtype: string75 - name: id76 dtype: string77 - name: content78 dtype: string79 splits:80 - name: train81 num_bytes: 4764927851.082 num_examples: 3272083 download_size: 125478619984 dataset_size: 4764927851.085- config_name: issues86 features:87 - name: repo_name88 dtype: string89 - name: content90 dtype: string91 - name: issue_id92 dtype: string93 splits:94 - name: train95 num_bytes: 31219575534.3848496 num_examples: 1554968297 download_size: 1648389904798 dataset_size: 31219575534.3848499- config_name: kaggle100 features:101 - name: content102 dtype: string103 - name: file_id104 dtype: string105 splits:106 - name: train107 num_bytes: 5228745262.0108 num_examples: 580195109 download_size: 2234440007110 dataset_size: 5228745262.0111- config_name: lhq112 features:113 - name: content114 dtype: string115 - name: metadata116 struct:117 - name: difficulty118 dtype: string119 - name: field120 dtype: string121 - name: topic122 dtype: string123 splits:124 - name: train125 num_bytes: 751273849.0126 num_examples: 7037500127 download_size: 272913202128 dataset_size: 751273849.0129- config_name: owm130 features:131 - name: url132 dtype: string133 - name: date134 dtype: timestamp[s]135 - name: metadata136 dtype: string137 - name: content138 dtype: string139 splits:140 - name: train141 num_bytes: 56294728333.0142 num_examples: 6315233143 download_size: 27160071916144 dataset_size: 56294728333.0145- config_name: stackoverflow146 features:147 - name: date148 dtype: string149 - name: nb_tokens150 dtype: int64151 - name: text_size152 dtype: int64153 - name: content154 dtype: string155 splits:156 - name: train157 num_bytes: 35548199612.0158 num_examples: 10404628159 download_size: 17008831030160 dataset_size: 35548199612.0161- config_name: wikipedia162 features:163 - name: content164 dtype: string165 - name: meta166 dtype: string167 - name: red_pajama_subset168 dtype: string169 splits:170 - name: train171 num_bytes: 21572720540.0172 num_examples: 6630651173 download_size: 12153445493174 dataset_size: 21572720540.0175configs:176- config_name: arxiv177 data_files:178 - split: train179 path: arxiv/train-*180- config_name: documentation181 data_files:182 - split: train183 path: documentation/train-*184- config_name: ir_cpp185 data_files:186 - split: train187 path: ir_cpp/train-*188- config_name: ir_low_resource189 data_files:190 - split: train191 path: ir_low_resource/train-*192- config_name: ir_python193 data_files:194 - split: train195 path: ir_python/train-*196- config_name: ir_rust197 data_files:198 - split: train199 path: ir_rust/train-*200- config_name: issues201 data_files:202 - split: train203 path: issues/train-*204- config_name: kaggle205 data_files:206 - split: train207 path: kaggle/train-*208- config_name: lhq209 data_files:210 - split: train211 path: lhq/train-*212- config_name: owm213 data_files:214 - split: train215 path: owm/train-*216- config_name: stackoverflow217 data_files:218 - split: train219 path: stackoverflow/train-*220- config_name: wikipedia221 data_files:222 - split: train223 path: wikipedia/train-*224---225 226# StarCoder2 Extras227 228This is the dataset of extra sources (besides Stack v2 code data) used to train the [StarCoder2](https://arxiv.org/abs/2402.19173) family of models. It contains the following subsets:229 230- Kaggle (`kaggle`): Kaggle notebooks from [Meta-Kaggle-Code](https://www.kaggle.com/datasets/kaggle/meta-kaggle-code) dataset, converted to scripts and prefixed with information on the Kaggle datasets used in the notebook. The file headers have a similar format to Jupyter Structured but the code content is only one single script.231- StackOverflow (`stackoverflow`): stackoverflow conversations from this [StackExchange dump](https://archive.org/details/stackexchange).232- Issues (`issues`): processed GitHub issues, same as the Stack v1 issues.233- OWM (`owm`): the [Open-Web-Math](https://huggingface.co/datasets/open-web-math/open-web-math) dataset.234- LHQ (`lhq`): Leandro's High quality dataset, it is a compilation of high quality code files from: APPS-train, CodeContests, GSM8K-train, GSM8K-SciRel, DeepMind-Mathematics, Rosetta-Code, MultiPL-T, ProofSteps, ProofSteps-lean.235- Wiki (`wikipedia`): the English subset of the Wikipedia dump in [RedPajama](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-1T).236- ArXiv (`arxiv`): the ArXiv subset of [RedPajama](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-1T) dataset, further processed the dataset only to retain latex source files and remove preambles, comments, macros, and bibliographies from these files.237- IR_language (`ir_cpp`, `ir_low_resource`, `ir_python`, `ir_rust`): these are intermediate representations of Python, Rust, C++ and other low resource languages.238- Documentation (`documentation`): documentation of popular libraries.239 240For more details on the processing of each subset, check the [StarCoder2 paper](https://arxiv.org/abs/2402.19173) or The Stack v2 [GitHub repository](https://github.com/bigcode-project/the-stack-v2/).241 242## Usage 243 244```python245from datasets import load_dataset246 247# replace `kaggle` with one of the config names listed above248ds = load_dataset("bigcode/starcoder2data-extras", "kaggle", split="train")249```250 251## Citation252 253```254@article{lozhkov2024starcoder,255 title={Starcoder 2 and the stack v2: The next generation},256 author={Lozhkov, Anton and Li, Raymond and Allal, Loubna Ben and Cassano, Federico and Lamy-Poirier, Joel and Tazi, Nouamane and Tang, Ao and Pykhtar, Dmytro and Liu, Jiawei and Wei, Yuxiang and others},257 journal={arXiv preprint arXiv:2402.19173},258 year={2024}259}260```