Team Ai
Datasetpublic

bigcode/starcoder2data-extras

StarCoder2 Extras This is the dataset of extra sources (besides Stack v2 code data) used to train the StarCoder2 family of models. It contains the following subsets: Kaggle (kaggle): Kaggle notebooks from Meta-Kaggle-Code dataset, converted to scripts and prefixed with information on the Kaggle datasets used in the notebook. The file headers have a similar format to Jupyter Structured but the code content is only one single script. StackOverflow (stackoverflow): stackoverflow… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/starcoder2data-extras.

sourceHugging Faceupdated 2y agoView on Hugging Face
13likes3.7kdownloads
README.md260 linesDownload Raw Back to root
1---2dataset_info:3- config_name: arxiv4  features:5  - name: content6    dtype: string7  splits:8  - name: train9    num_bytes: 89223183645.010    num_examples: 155830611  download_size: 4091118687612  dataset_size: 89223183645.013- config_name: documentation14  features:15  - name: project16    dtype: string17  - name: source18    dtype: string19  - name: language20    dtype: string21  - name: content22    dtype: string23  splits:24  - name: train25    num_bytes: 5421472234.026    num_examples: 5973327  download_size: 185345192228  dataset_size: 5421472234.029- config_name: ir_cpp30  features:31  - name: __index_level_0__32    dtype: string33  - name: id34    dtype: string35  - name: content36    dtype: string37  splits:38  - name: train39    num_bytes: 102081135272.040    num_examples: 291665541  download_size: 2604797842242  dataset_size: 102081135272.043- config_name: ir_low_resource44  features:45  - name: __index_level_0__46    dtype: string47  - name: id48    dtype: string49  - name: content50    dtype: string51  - name: size52    dtype: int6453  splits:54  - name: train55    num_bytes: 10383382043.056    num_examples: 39398857  download_size: 246451360358  dataset_size: 10383382043.059- config_name: ir_python60  features:61  - name: id62    dtype: string63  - name: content64    dtype: string65  splits:66  - name: train67    num_bytes: 12446664464.068    num_examples: 15450769  download_size: 303929762570  dataset_size: 12446664464.071- config_name: ir_rust72  features:73  - name: __index_level_0__74    dtype: string75  - name: id76    dtype: string77  - name: content78    dtype: string79  splits:80  - name: train81    num_bytes: 4764927851.082    num_examples: 3272083  download_size: 125478619984  dataset_size: 4764927851.085- config_name: issues86  features:87  - name: repo_name88    dtype: string89  - name: content90    dtype: string91  - name: issue_id92    dtype: string93  splits:94  - name: train95    num_bytes: 31219575534.3848496    num_examples: 1554968297  download_size: 1648389904798  dataset_size: 31219575534.3848499- config_name: kaggle100  features:101  - name: content102    dtype: string103  - name: file_id104    dtype: string105  splits:106  - name: train107    num_bytes: 5228745262.0108    num_examples: 580195109  download_size: 2234440007110  dataset_size: 5228745262.0111- config_name: lhq112  features:113  - name: content114    dtype: string115  - name: metadata116    struct:117    - name: difficulty118      dtype: string119    - name: field120      dtype: string121    - name: topic122      dtype: string123  splits:124  - name: train125    num_bytes: 751273849.0126    num_examples: 7037500127  download_size: 272913202128  dataset_size: 751273849.0129- config_name: owm130  features:131  - name: url132    dtype: string133  - name: date134    dtype: timestamp[s]135  - name: metadata136    dtype: string137  - name: content138    dtype: string139  splits:140  - name: train141    num_bytes: 56294728333.0142    num_examples: 6315233143  download_size: 27160071916144  dataset_size: 56294728333.0145- config_name: stackoverflow146  features:147  - name: date148    dtype: string149  - name: nb_tokens150    dtype: int64151  - name: text_size152    dtype: int64153  - name: content154    dtype: string155  splits:156  - name: train157    num_bytes: 35548199612.0158    num_examples: 10404628159  download_size: 17008831030160  dataset_size: 35548199612.0161- config_name: wikipedia162  features:163  - name: content164    dtype: string165  - name: meta166    dtype: string167  - name: red_pajama_subset168    dtype: string169  splits:170  - name: train171    num_bytes: 21572720540.0172    num_examples: 6630651173  download_size: 12153445493174  dataset_size: 21572720540.0175configs:176- config_name: arxiv177  data_files:178  - split: train179    path: arxiv/train-*180- config_name: documentation181  data_files:182  - split: train183    path: documentation/train-*184- config_name: ir_cpp185  data_files:186  - split: train187    path: ir_cpp/train-*188- config_name: ir_low_resource189  data_files:190  - split: train191    path: ir_low_resource/train-*192- config_name: ir_python193  data_files:194  - split: train195    path: ir_python/train-*196- config_name: ir_rust197  data_files:198  - split: train199    path: ir_rust/train-*200- config_name: issues201  data_files:202  - split: train203    path: issues/train-*204- config_name: kaggle205  data_files:206  - split: train207    path: kaggle/train-*208- config_name: lhq209  data_files:210  - split: train211    path: lhq/train-*212- config_name: owm213  data_files:214  - split: train215    path: owm/train-*216- config_name: stackoverflow217  data_files:218  - split: train219    path: stackoverflow/train-*220- config_name: wikipedia221  data_files:222  - split: train223    path: wikipedia/train-*224---225 226# StarCoder2 Extras227 228This is the dataset of extra sources (besides Stack v2 code data) used to train the [StarCoder2](https://arxiv.org/abs/2402.19173) family of models. It contains the following subsets:229 230- Kaggle (`kaggle`): Kaggle notebooks from [Meta-Kaggle-Code](https://www.kaggle.com/datasets/kaggle/meta-kaggle-code) dataset, converted to scripts and prefixed with information on the Kaggle datasets used in the notebook. The file headers have a similar format to Jupyter Structured but the code content is only one single script.231- StackOverflow (`stackoverflow`): stackoverflow conversations from this [StackExchange dump](https://archive.org/details/stackexchange).232- Issues (`issues`): processed GitHub issues, same as the Stack v1 issues.233- OWM (`owm`): the [Open-Web-Math](https://huggingface.co/datasets/open-web-math/open-web-math) dataset.234- LHQ (`lhq`): Leandro's High quality dataset, it is a compilation of high quality code files from: APPS-train, CodeContests, GSM8K-train, GSM8K-SciRel, DeepMind-Mathematics, Rosetta-Code, MultiPL-T, ProofSteps, ProofSteps-lean.235- Wiki (`wikipedia`): the English subset of the Wikipedia dump in [RedPajama](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-1T).236- ArXiv (`arxiv`): the ArXiv subset of [RedPajama](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-1T) dataset, further processed the dataset only to retain latex source files and remove preambles, comments, macros, and bibliographies from these files.237- IR_language (`ir_cpp`, `ir_low_resource`, `ir_python`, `ir_rust`): these are intermediate representations of Python, Rust, C++ and other low resource languages.238- Documentation (`documentation`): documentation of popular libraries.239 240For more details on the processing of each subset, check the [StarCoder2 paper](https://arxiv.org/abs/2402.19173) or The Stack v2 [GitHub repository](https://github.com/bigcode-project/the-stack-v2/).241 242## Usage 243 244```python245from datasets import load_dataset246 247# replace `kaggle` with one of the config names listed above248ds = load_dataset("bigcode/starcoder2data-extras", "kaggle", split="train")249```250 251## Citation252 253```254@article{lozhkov2024starcoder,255  title={Starcoder 2 and the stack v2: The next generation},256  author={Lozhkov, Anton and Li, Raymond and Allal, Loubna Ben and Cassano, Federico and Lamy-Poirier, Joel and Tazi, Nouamane and Tang, Ao and Pykhtar, Dmytro and Liu, Jiawei and Wei, Yuxiang and others},257  journal={arXiv preprint arXiv:2402.19173},258  year={2024}259}260```