bleugreen/typescript-chunks
typescript-chunks A dataset of TypeScript snippets, processed from the typescript subset of the-stack-smol. Processing Each source file is parsed with the TypeScript AST and queried for 'semantic chunks' of the following types. FunctionDeclaration ---- 8205 ArrowFunction --------- 33890 ClassDeclaration ------- 5325 InterfaceDeclaration -- 12884 EnumDeclaration --------- 518 TypeAliasDeclaration --- 3580 MethodDeclaration ----- 24713 Leading comments are added… See the full description on the dataset page: https://huggingface.co/datasets/bleugreen/typescript-chunks.
334
1from datasets import Dataset, load_dataset2from transformers import AutoTokenizer3 4tokenizer = AutoTokenizer.from_pretrained('models/RedPajama-INCITE-Instruct-7B')5max_seq = 20486 7def make_prompt(code):8 return f'Below is an instruction that describes a task. Write a response that appropriately completes the request.\n\n### Instruction:\n{code}\n\n### Response:\n'9 10 11def is_not_too_long(data):12 encoded = tokenizer.encode(make_prompt(data['content']))13 return len(encoded) < max_seq14 15def deduplicate_dicts(dicts):16 seen = {}17 result = []18 for d in dicts:19 content = d.get('content')20 if content not in seen:21 seen[content] = True22 result.append(d)23 return result24 25dataset = load_dataset('json', data_files='ts_parser/ts-chunks.jsonl')26 27data_short = dataset.filter(is_not_too_long)28 29dedup = deduplicate_dicts(data_short['train'])30 31data_short_dedup = Dataset.from_list(dedup)32print(data_short_dedup)33 34data_short_dedup.to_json('typescript-chunks.json')35 