Team Ai
Datasetpublic

bleugreen/typescript-chunks

typescript-chunks A dataset of TypeScript snippets, processed from the typescript subset of the-stack-smol. Processing Each source file is parsed with the TypeScript AST and queried for 'semantic chunks' of the following types. FunctionDeclaration ---- 8205 ArrowFunction --------- 33890 ClassDeclaration ------- 5325 InterfaceDeclaration -- 12884 EnumDeclaration --------- 518 TypeAliasDeclaration --- 3580 MethodDeclaration ----- 24713 Leading comments are added… See the full description on the dataset page: https://huggingface.co/datasets/bleugreen/typescript-chunks.

sourceHugging Faceupdated 3y agoView on Hugging Face
3likes34downloads
process.py35 linesDownload Raw Back to root
1from datasets import Dataset, load_dataset2from transformers import AutoTokenizer3 4tokenizer = AutoTokenizer.from_pretrained('models/RedPajama-INCITE-Instruct-7B')5max_seq = 20486 7def make_prompt(code):8    return f'Below is an instruction that describes a task. Write a response that appropriately completes the request.\n\n### Instruction:\n{code}\n\n### Response:\n'9 10 11def is_not_too_long(data):12    encoded = tokenizer.encode(make_prompt(data['content']))13    return len(encoded) < max_seq14 15def deduplicate_dicts(dicts):16    seen = {}17    result = []18    for d in dicts:19        content = d.get('content')20        if content not in seen:21            seen[content] = True22            result.append(d)23    return result24 25dataset = load_dataset('json', data_files='ts_parser/ts-chunks.jsonl')26 27data_short = dataset.filter(is_not_too_long)28 29dedup = deduplicate_dicts(data_short['train'])30 31data_short_dedup = Dataset.from_list(dedup)32print(data_short_dedup)33 34data_short_dedup.to_json('typescript-chunks.json')35