Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ajibawa-2023 /JavaScript-Code-LargeJavaScript-Code-Large JavaScript-Code-Large is a large-scale corpus of JavaScript source code comprising around 5 million JavaScript files. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the JavaScript ecosystem. By providing a high-volume, language-specific corpus, JavaScript-Code-Large enables systematic experimentation in JavaScript-focused model training, domain adaptation… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/JavaScript-Code-Large.texttext-generation1M<n<10M34 likes2.5k downloads8mo agoHugging Face02angie-chen55 /javascript-github-codetext10M<n<100M18 likes2.1k downloads4y agoHugging Face03thedruid831 /JavaScript-Code-LargeJavaScript-Code-Large JavaScript-Code-Large is a large-scale corpus of JavaScript source code comprising around 5 million JavaScript files. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the JavaScript ecosystem. By providing a high-volume, language-specific corpus, JavaScript-Code-Large enables systematic experimentation in JavaScript-focused model training, domain adaptation… See the full description on the dataset page: https://huggingface.co/datasets/thedruid831/JavaScript-Code-Large.texttext-generation1M<n<10M0 likes561 downloads8mo agoHugging Face04liuhangbiao /JavaScript-Code-LargeJavaScript-Code-Large JavaScript-Code-Large is a large-scale corpus of JavaScript source code comprising around 5 million JavaScript files. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the JavaScript ecosystem. By providing a high-volume, language-specific corpus, JavaScript-Code-Large enables systematic experimentation in JavaScript-focused model training, domain adaptation… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/JavaScript-Code-Large.texttext-generation1M<n<10M0 likes507 downloads6mo agoHugging Face05Ujjwal-Tyagi /JavaScript-Code-LargeJavaScript-Code-Large JavaScript-Code-Large is a large-scale corpus of JavaScript source code comprising around 5 million JavaScript files. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the JavaScript ecosystem. By providing a high-volume, language-specific corpus, JavaScript-Code-Large enables systematic experimentation in JavaScript-focused model training, domain adaptation… See the full description on the dataset page: https://huggingface.co/datasets/Ujjwal-Tyagi/JavaScript-Code-Large.texttext-generation1M<n<10M1 likes484 downloads6mo agoHugging Face06hchautran /javascript-mediumtext100K<n<1M2 likes406 downloads4y agoHugging Face07nomic-ai /cornstack-javascript-v1 CoRNStack Javascript Dataset The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code, CodeRankEmbed, and CodeRankLLM. CoRNStack Dataset Curation Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-javascript-v1.text1M<n<10M5 likes364 downloads2y agoHugging Face08hchautran /javascripttext100K<n<1M0 likes270 downloads4y agoHugging Face09ThomasTheMaker /arc-stack-javascripttabular10M<n<100M0 likes253 downloads11mo agoHugging Face10Mo7art /Stack2Graph_KG_javascript Javascript StackOverflow Knowledge Graph Summary This Hugging Face dataset repository contains the Javascript shard of the Stack2Graph StackOverflow Knowledge Graph. Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper. The artifact is optimized for graph-based retrieval, SPARQL analytics, and retrieval-augmented question answering over Stack Overflow content.… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_KG_javascript.100M<n<1B0 likes250 downloads3mo agoHugging Face11hchautran /javascript-smalltext100K<n<1M4 likes192 downloads4y agoHugging Face12verify-ppt /marin-starcoderdata_javascript0 likes169 downloads6mo agoHugging Face13corniclr25 /stack-mined-javascript-v1text1M<n<10M0 likes125 downloads2y agoHugging Face14CM /codexglue_code2text_javascript Dataset Card for "codexglue_code2text_javascript" More Information needed text10K<n<100K12 likes123 downloads3y agoHugging Face15hongliu9903 /stack_edu_javascripttabular10M<n<100M0 likes114 downloads1y agoHugging Face16Nan-Do /code-search-net-javascript Dataset Card for "code-search-net-javascript" Dataset Summary This dataset is the JavaScript portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the function does. Languages The dataset's comments are in English and the functions are coded in JavaScript Data Splits Train, test, validation labels are… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-javascript.texttext-generation100K<n<1M7 likes109 downloads3y agoHugging Face17h4iku /coconut_javascript2010_preprocessedtext100K<n<1M0 likes102 downloads4y agoHugging Face18CoIR-Retrieval /CodeSearchNet-ccr-javascript-queries-corpus Dataset Card for "CodeSearchNet-ccr-javascript-queries-corpus" More Information needed text100K<n<1M0 likes99 downloads2y agoHugging Face19ars-1 /autotrain-data-javascript-traing-1 AutoTrain Dataset for project: javascript-traing-1 Dataset Description This dataset has been automatically processed by AutoTrain for project javascript-traing-1. Languages The BCP-47 code for the dataset's language is unk. Dataset Structure Data Instances A sample from this dataset looks as follows: [ { "target": "test/NavbarSpec.js", "feat_repo_name": "aabenoja/react-bootstrap", "text": "import React from 'react';\nimport… See the full description on the dataset page: https://huggingface.co/datasets/ars-1/autotrain-data-javascript-traing-1.textsummarization100K<n<1M0 likes96 downloads3y agoHugging Face20AmareshHebbar /leetcode-codegen-javascript LeetCode Code-Gen Dataset — JavaScript 631 rows. Given a problem statement, its input/output examples, and a required algorithm/technique, generate a correct JavaScript solution. Part of a 4-language collection built from the same source: see the sibling Python, Java, C++, and JavaScript datasets. Verification Not execution-verified. There is currently no compiler/runtime harness for this language in the build pipeline (only Python has one). Rows are extracted… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-javascript.texttext-generationn<1K1 likes93 downloads3mo agoHugging Face21CoIR-Retrieval /CodeSearchNet-javascript-qrels Dataset Card for "CodeSearchNet-javascript-qrels" More Information needed text10K<n<100K0 likes89 downloads2y agoHugging Face22CoIR-Retrieval /CodeSearchNet-javascript-queries-corpus Dataset Card for "CodeSearchNet-javascript-queries-corpus" More Information needed text100K<n<1M1 likes75 downloads2y agoHugging Face23nthakur /cornstack-javascript-v1-tevatron-1Mtext100K<n<1M0 likes71 downloads1y agoHugging Face24Mo7art /Stack2Graph_VD_javascript Javascript StackOverflow Vector Dataset Summary This Hugging Face dataset repository contains the Javascript shard of the Stack2Graph vector-database component as restorable Qdrant artifacts plus portable Parquet fallback files. Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper. The artifacts are intended for semantic and hybrid retrieval, graph entry-point finding… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_VD_javascript.feature-extraction0 likes66 downloads3mo agoHugging Face25h4iku /coconut_javascript2010 Dataset Card for CoCoNuT-JavaScript(2010) Dataset Summary Part of the data used to train the models in the "CoCoNuT: Combining Context-Aware Neural Translation Models using Ensemble for Program Repair" paper. These datasets contain raw data extracted from GitHub, GitLab, and Bitbucket, and have neither been shuffled nor tokenized. The year in the dataset’s name is the cutting year that shows the year of the newest commit in the dataset. Languages JavaScript… See the full description on the dataset page: https://huggingface.co/datasets/h4iku/coconut_javascript2010.text1M<n<10M2 likes62 downloads3y agoHugging Face26saurabh5 /rlvr-code-data-JavaScripttext100K<n<1M1 likes61 downloads1y agoHugging Face27semeru /code-text-javascript Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-text/javascript in Semeru CodeXGLUE -- Code-To-Text Task Definition The task is to generate natural language comments for a code, and evaluted by smoothed bleu-4 score. Dataset The dataset we use comes from CodeSearchNet and we filter the dataset as the following:… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-text-javascript.text10K<n<100K10 likes60 downloads4y agoHugging Face28CoIR-Retrieval /CodeSearchNet-ccr-javascript-qrels Dataset Card for "CodeSearchNet-ccr-javascript-qrels" More Information needed text10K<n<100K0 likes59 downloads2y agoHugging Face29finbarr /rlvr-code-data-javascript-editedtext100K<n<1M1 likes43 downloads1y agoHugging Face30Shuu12121 /javascript-treesitter-dedupe-filtered-datasetsV2 Javascript CodeSearch Dataset (Shuu12121/javascript-treesitter-dedupe-filtered-datasetsV2) Dataset Description This dataset contains JavaScript functions and methods paired with their JSDoc comments, extracted from open-source JavaScript repositories on GitHub. It is formatted similarly to the CodeSearchNet challenge dataset. Each entry includes: code: The source code of a javascript function or method. docstring: The docstring or Javadoc associated with the… See the full description on the dataset page: https://huggingface.co/datasets/Shuu12121/javascript-treesitter-dedupe-filtered-datasetsV2.text100K<n<1M0 likes43 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.