bleugreen/typescript-chunks
typescript-chunks A dataset of TypeScript snippets, processed from the typescript subset of the-stack-smol. Processing Each source file is parsed with the TypeScript AST and queried for 'semantic chunks' of the following types. FunctionDeclaration ---- 8205 ArrowFunction --------- 33890 ClassDeclaration ------- 5325 InterfaceDeclaration -- 12884 EnumDeclaration --------- 518 TypeAliasDeclaration --- 3580 MethodDeclaration ----- 24713 Leading comments are added… See the full description on the dataset page: https://huggingface.co/datasets/bleugreen/typescript-chunks.
334
1---2task_categories:3- text-classification4- text2text-generation5- summarization6language:7- en8---9 10# typescript-chunks11A dataset of TypeScript snippets, processed from the typescript subset of [the-stack-smol](https://huggingface.co/datasets/bigcode/the-stack-smol).12 13 14# Processing15- Each source file is parsed with the TypeScript AST and queried for 'semantic chunks' of the following types. 16```17FunctionDeclaration ---- 820518ArrowFunction --------- 3389019ClassDeclaration ------- 532520InterfaceDeclaration -- 1288421EnumDeclaration --------- 51822TypeAliasDeclaration --- 358023MethodDeclaration ----- 2471324```25- Leading comments are added to the front of `content`26- Removed all chunks over max sequence length (2048)27- Deduplicated / cleaned up28- Generated instructions / summaries with `gpt-3.5-turbo` (in progress)29 30 31 32# Dataset Structure33```python34from datasets import load_dataset35load_dataset("bleugreen/typescript-chunks")36 37DatasetDict({38 train: Dataset({39 features: ['type', 'content', 'repo', 'path', 'language'],40 num_rows: 8911541 })42})43```