datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cross_code_eval_typescripttypescript-codegithub-file-programs-dataset-typescriptdev_set_v2_a1_crosscodeeval_typescript_20260805_125428a1-crosscodeeval_typescript-tb2-leonardo-20260731marin-starcoderdata_typescriptstack_edu_typescripttypescript-instruct-20kWhy always Python?
I get 20,000 TypeScript code from The Stack and generate {"instruction", "output"} pairs (based on gpt-3.5-turbo)
Using this dataset for finetune code generation model just for TypeScript
Make web developers great again !
a1_crosscodeeval_typescripttypescript-treesitter-dedupe-filtered-datasetsV2
Typescript CodeSearch Dataset (Shuu12121/typescript-treesitter-dedupe-filtered-datasetsV2)
Dataset Description
This dataset contains TypeScript functions and methods paired with their TSDoc comments, extracted from open-source TypeScript repositories on GitHub.
It is formatted similarly to the CodeSearchNet challenge dataset.
Each entry includes:
code: The source code of a typescript function or method.
docstring: The docstring or Javadoc associated with the… See the full description on the dataset page: https://huggingface.co/datasets/Shuu12121/typescript-treesitter-dedupe-filtered-datasetsV2.typescript-treesitter-filtered-datasetsV2
Typescript CodeSearch Dataset (Shuu12121/typescript-treesitter-filtered-datasetsV2)
Dataset Description
This dataset contains TypeScript functions and methods paired with their TSDoc comments, extracted from open-source TypeScript repositories on GitHub.
It is formatted similarly to the CodeSearchNet challenge dataset.
Each entry includes:
code: The source code of a typescript function or method.
docstring: The docstring or Javadoc associated with the function/method.… See the full description on the dataset page: https://huggingface.co/datasets/Shuu12121/typescript-treesitter-filtered-datasetsV2.exp_rpt_crosscodeeval-typescript_10kexp_rpt_crosscodeeval-typescriptswebench_verified_random_100_folders_a1_crosscodeeval_typescript_20260324_072751typescript-instruct
typescript-instruct
A dataset of TypeScript snippets, processed from the typescript subset of the-stack-smol.
Processing
Each source file is parsed with the TypeScript AST and queried for 'semantic chunks' of the following types.
ClassDeclaration - 2401
ArrowFunction - 16443… See the full description on the dataset page: https://huggingface.co/datasets/bleugreen/typescript-instruct.smollm3-stack-v2-TypeScripttypescript-dataset
TypeScript Advanced Reasoning Dataset
This dataset provides a large collection of advanced TypeScript reasoning tasks designed to train models that understand and operate within the TypeScript type system at an expert level. The content focuses on type theory, generic inference, discriminated unions, template literal behavior, narrowing rules, static analysis, and complex type transformations.
Each entry is formatted as a compact JSONL instruction output pair so it can be… See the full description on the dataset page: https://huggingface.co/datasets/grenishrai/typescript-dataset.terminal_bench_2_a1_crosscodeeval_typescript_20260324_025704terminal_bench_2_a1_crosscodeeval_typescript_20260325_000759swebench_verified_random_100_folders_a1_crosscodeeval_typescript_20260625_215021typescript-chunks
typescript-chunks
A dataset of TypeScript snippets, processed from the typescript subset of the-stack-smol.
Processing
Each source file is parsed with the TypeScript AST and queried for 'semantic chunks' of the following types.
FunctionDeclaration ---- 8205
ArrowFunction --------- 33890
ClassDeclaration ------- 5325
InterfaceDeclaration -- 12884
EnumDeclaration --------- 518
TypeAliasDeclaration --- 3580
MethodDeclaration ----- 24713
Leading comments are added to the… See the full description on the dataset page: https://huggingface.co/datasets/bleugreen/typescript-chunks.refineid-stack-v3-typescript
RefineID-format TypeScript identifier recovery from The Stack v3
This public benchmark contains 1,000 examples from 1,000
source repositories and 3,029 masked identifier positions. Each
example is a full source file with one chosen identifier name masked at every
eligible identifier token position. The target is the original name.
Format and use
test.csv has no header and exactly three columns in the same order as the
original RefineID benchmark: id, code_masked… See the full description on the dataset page: https://huggingface.co/datasets/D4vidHuang/refineid-stack-v3-typescript.dev_set_v2_a1_crosscodeeval_typescript_20260324_201725swebench_verified_random_100_folders_a1_crosscodeeval_typescript_20260325_072852rlvr-code-data-typescript-editedtypescript_ins_binarized
Dataset Card for "typescript_ins_binarized"
More Information needed
typescript-instruct-20k-v2cWhy always Python?
I get 20,000 TypeScript code from The Stack and generate {"instruction", "output"} pairs (based on gpt-3.5-turbo)
Using this dataset for finetune code generation model just for TypeScript
Make web developers great again !
typescript-testtypescript-jestStack2Graph_VD_typescript
Typescript StackOverflow Vector Dataset
Summary
This Hugging Face dataset repository contains the Typescript shard of the Stack2Graph vector-database component as restorable Qdrant artifacts plus portable Parquet fallback files.
Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper.
The artifacts are intended for semantic and hybrid retrieval, graph entry-point finding… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_VD_typescript.
