Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ZHENGRAN /cross_code_eval_typescripttext1K<n<10K1 likes600 downloads2y agoHugging Face02petrpan26 /typescript-codetabular100K<n<1M3 likes210 downloads3y agoHugging Face03Shuu12121 /github-file-programs-dataset-typescripttext1M<n<10M1 likes151 downloads9mo agoHugging Face04laion /dev_set_v2_a1_crosscodeeval_typescript_20260805_125428text1K<n<10K0 likes128 downloads2mo agoHugging Face05laion /a1-crosscodeeval_typescript-tb2-leonardo-20260731text1K<n<10K1 likes122 downloads2mo agoHugging Face06verify-ppt /marin-starcoderdata_typescript0 likes108 downloads6mo agoHugging Face07hongliu9903 /stack_edu_typescripttabular1M<n<10M0 likes96 downloads1y agoHugging Face08mhhmm /typescript-instruct-20kWhy always Python? I get 20,000 TypeScript code from The Stack and generate {"instruction", "output"} pairs (based on gpt-3.5-turbo) Using this dataset for finetune code generation model just for TypeScript Make web developers great again ! texttext-generation10K<n<100K11 likes86 downloads3y agoHugging Face09DCAgent /a1_crosscodeeval_typescripttext10K<n<100K0 likes83 downloads7mo agoHugging Face10Shuu12121 /typescript-treesitter-dedupe-filtered-datasetsV2 Typescript CodeSearch Dataset (Shuu12121/typescript-treesitter-dedupe-filtered-datasetsV2) Dataset Description This dataset contains TypeScript functions and methods paired with their TSDoc comments, extracted from open-source TypeScript repositories on GitHub. It is formatted similarly to the CodeSearchNet challenge dataset. Each entry includes: code: The source code of a typescript function or method. docstring: The docstring or Javadoc associated with the… See the full description on the dataset page: https://huggingface.co/datasets/Shuu12121/typescript-treesitter-dedupe-filtered-datasetsV2.text100K<n<1M0 likes72 downloads1y agoHugging Face11Shuu12121 /typescript-treesitter-filtered-datasetsV2 Typescript CodeSearch Dataset (Shuu12121/typescript-treesitter-filtered-datasetsV2) Dataset Description This dataset contains TypeScript functions and methods paired with their TSDoc comments, extracted from open-source TypeScript repositories on GitHub. It is formatted similarly to the CodeSearchNet challenge dataset. Each entry includes: code: The source code of a typescript function or method. docstring: The docstring or Javadoc associated with the function/method.… See the full description on the dataset page: https://huggingface.co/datasets/Shuu12121/typescript-treesitter-filtered-datasetsV2.text100K<n<1M2 likes69 downloads1y agoHugging Face12DCAgent /exp_rpt_crosscodeeval-typescript_10ktext10K<n<100K0 likes66 downloads7mo agoHugging Face13DCAgent /exp_rpt_crosscodeeval-typescripttext1K<n<10K0 likes61 downloads8mo agoHugging Face14DCAgent2 /swebench_verified_random_100_folders_a1_crosscodeeval_typescript_20260324_072751textn<1K0 likes55 downloads7mo agoHugging Face15bleugreen /typescript-instruct typescript-instruct A dataset of TypeScript snippets, processed from the typescript subset of the-stack-smol. Processing Each source file is parsed with the TypeScript AST and queried for 'semantic chunks' of the following types. ClassDeclaration - 2401 ArrowFunction - 16443… See the full description on the dataset page: https://huggingface.co/datasets/bleugreen/typescript-instruct.texttext-classification10K<n<100K8 likes54 downloads3y agoHugging Face16verify-ppt /smollm3-stack-v2-TypeScript0 likes49 downloads7mo agoHugging Face17grenishrai /typescript-dataset TypeScript Advanced Reasoning Dataset This dataset provides a large collection of advanced TypeScript reasoning tasks designed to train models that understand and operate within the TypeScript type system at an expert level. The content focuses on type theory, generic inference, discriminated unions, template literal behavior, narrowing rules, static analysis, and complex type transformations. Each entry is formatted as a compact JSONL instruction output pair so it can be… See the full description on the dataset page: https://huggingface.co/datasets/grenishrai/typescript-dataset.texttext-generation1K<n<10K1 likes48 downloads10mo agoHugging Face18DCAgent2 /terminal_bench_2_a1_crosscodeeval_typescript_20260324_025704textn<1K0 likes44 downloads7mo agoHugging Face19DCAgent2 /terminal_bench_2_a1_crosscodeeval_typescript_20260325_000759textn<1K0 likes44 downloads7mo agoHugging Face20DCAgent2 /swebench_verified_random_100_folders_a1_crosscodeeval_typescript_20260625_215021text1K<n<10K0 likes42 downloads3mo agoHugging Face21bleugreen /typescript-chunks typescript-chunks A dataset of TypeScript snippets, processed from the typescript subset of the-stack-smol. Processing Each source file is parsed with the TypeScript AST and queried for 'semantic chunks' of the following types. FunctionDeclaration ---- 8205 ArrowFunction --------- 33890 ClassDeclaration ------- 5325 InterfaceDeclaration -- 12884 EnumDeclaration --------- 518 TypeAliasDeclaration --- 3580 MethodDeclaration ----- 24713 Leading comments are added to the… See the full description on the dataset page: https://huggingface.co/datasets/bleugreen/typescript-chunks.texttext-classification10K<n<100K3 likes41 downloads3y agoHugging Face22D4vidHuang /refineid-stack-v3-typescript RefineID-format TypeScript identifier recovery from The Stack v3 This public benchmark contains 1,000 examples from 1,000 source repositories and 3,029 masked identifier positions. Each example is a full source file with one chosen identifier name masked at every eligible identifier token position. The target is the original name. Format and use test.csv has no header and exactly three columns in the same order as the original RefineID benchmark: id, code_masked… See the full description on the dataset page: https://huggingface.co/datasets/D4vidHuang/refineid-stack-v3-typescript.text1K<n<10K0 likes40 downloads12d agoHugging Face23DCAgent2 /dev_set_v2_a1_crosscodeeval_typescript_20260324_201725textn<1K0 likes36 downloads7mo agoHugging Face24DCAgent2 /swebench_verified_random_100_folders_a1_crosscodeeval_typescript_20260325_072852textn<1K0 likes30 downloads7mo agoHugging Face25finbarr /rlvr-code-data-typescript-editedtext100K<n<1M0 likes27 downloads1y agoHugging Face26jan-hq /typescript_ins_binarized Dataset Card for "typescript_ins_binarized" More Information needed text10K<n<100K0 likes26 downloads3y agoHugging Face27mhhmm /typescript-instruct-20k-v2cWhy always Python? I get 20,000 TypeScript code from The Stack and generate {"instruction", "output"} pairs (based on gpt-3.5-turbo) Using this dataset for finetune code generation model just for TypeScript Make web developers great again ! texttext-generation10K<n<100K5 likes26 downloads3y agoHugging Face28petrpan26 /typescript-testtabular10K<n<100K0 likes25 downloads3y agoHugging Face29petrpan26 /typescript-jesttabular10K<n<100K0 likes19 downloads3y agoHugging Face30Mo7art /Stack2Graph_VD_typescript Typescript StackOverflow Vector Dataset Summary This Hugging Face dataset repository contains the Typescript shard of the Stack2Graph vector-database component as restorable Qdrant artifacts plus portable Parquet fallback files. Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper. The artifacts are intended for semantic and hybrid retrieval, graph entry-point finding… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_VD_typescript.feature-extraction0 likes19 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.