Team Ai
Datasetpublic

bleugreen/typescript-chunks

typescript-chunks A dataset of TypeScript snippets, processed from the typescript subset of the-stack-smol. Processing Each source file is parsed with the TypeScript AST and queried for 'semantic chunks' of the following types. FunctionDeclaration ---- 8205 ArrowFunction --------- 33890 ClassDeclaration ------- 5325 InterfaceDeclaration -- 12884 EnumDeclaration --------- 518 TypeAliasDeclaration --- 3580 MethodDeclaration ----- 24713 Leading comments are added… See the full description on the dataset page: https://huggingface.co/datasets/bleugreen/typescript-chunks.

sourceHugging Faceupdated 3y agoView on Hugging Face
3likes34downloads
README.md43 linesDownload Raw Back to root
1---2task_categories:3- text-classification4- text2text-generation5- summarization6language:7- en8---9 10# typescript-chunks11A dataset of TypeScript snippets, processed from the typescript subset of [the-stack-smol](https://huggingface.co/datasets/bigcode/the-stack-smol).12 13 14# Processing15- Each source file is parsed with the TypeScript AST and queried for 'semantic chunks' of the following types. 16```17FunctionDeclaration ---- 820518ArrowFunction --------- 3389019ClassDeclaration ------- 532520InterfaceDeclaration -- 1288421EnumDeclaration --------- 51822TypeAliasDeclaration --- 358023MethodDeclaration ----- 2471324```25- Leading comments are added to the front of `content`26- Removed all chunks over max sequence length (2048)27- Deduplicated / cleaned up28- Generated instructions / summaries with `gpt-3.5-turbo` (in progress)29 30 31 32# Dataset Structure33```python34from datasets import load_dataset35load_dataset("bleugreen/typescript-chunks")36 37DatasetDict({38    train: Dataset({39        features: ['type', 'content', 'repo', 'path', 'language'],40        num_rows: 8911541    })42})43```