Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ajibawa-2023 /JavaScript-Code-LargeJavaScript-Code-Large JavaScript-Code-Large is a large-scale corpus of JavaScript source code comprising around 5 million JavaScript files. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the JavaScript ecosystem. By providing a high-volume, language-specific corpus, JavaScript-Code-Large enables systematic experimentation in JavaScript-focused model training, domain adaptation… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/JavaScript-Code-Large.texttext-generation1M<n<10M34 likes2.5k downloads8mo agoHugging Face02angie-chen55 /javascript-github-codetext10M<n<100M18 likes2.1k downloads4y agoHugging Face03thedruid831 /JavaScript-Code-LargeJavaScript-Code-Large JavaScript-Code-Large is a large-scale corpus of JavaScript source code comprising around 5 million JavaScript files. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the JavaScript ecosystem. By providing a high-volume, language-specific corpus, JavaScript-Code-Large enables systematic experimentation in JavaScript-focused model training, domain adaptation… See the full description on the dataset page: https://huggingface.co/datasets/thedruid831/JavaScript-Code-Large.texttext-generation1M<n<10M0 likes561 downloads8mo agoHugging Face04liuhangbiao /JavaScript-Code-LargeJavaScript-Code-Large JavaScript-Code-Large is a large-scale corpus of JavaScript source code comprising around 5 million JavaScript files. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the JavaScript ecosystem. By providing a high-volume, language-specific corpus, JavaScript-Code-Large enables systematic experimentation in JavaScript-focused model training, domain adaptation… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/JavaScript-Code-Large.texttext-generation1M<n<10M0 likes507 downloads6mo agoHugging Face05Ujjwal-Tyagi /JavaScript-Code-LargeJavaScript-Code-Large JavaScript-Code-Large is a large-scale corpus of JavaScript source code comprising around 5 million JavaScript files. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the JavaScript ecosystem. By providing a high-volume, language-specific corpus, JavaScript-Code-Large enables systematic experimentation in JavaScript-focused model training, domain adaptation… See the full description on the dataset page: https://huggingface.co/datasets/Ujjwal-Tyagi/JavaScript-Code-Large.texttext-generation1M<n<10M1 likes484 downloads6mo agoHugging Face06hchautran /javascript-mediumtext100K<n<1M2 likes406 downloads4y agoHugging Face07nomic-ai /cornstack-javascript-v1 CoRNStack Javascript Dataset The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code, CodeRankEmbed, and CodeRankLLM. CoRNStack Dataset Curation Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-javascript-v1.text1M<n<10M5 likes364 downloads2y agoHugging Face08hchautran /javascripttext100K<n<1M0 likes270 downloads4y agoHugging Face09ThomasTheMaker /arc-stack-javascripttabular10M<n<100M0 likes253 downloads11mo agoHugging Face10hchautran /javascript-smalltext100K<n<1M4 likes192 downloads4y agoHugging Face11corniclr25 /stack-mined-javascript-v1text1M<n<10M0 likes125 downloads2y agoHugging Face12CM /codexglue_code2text_javascript Dataset Card for "codexglue_code2text_javascript" More Information needed text10K<n<100K12 likes123 downloads3y agoHugging Face13hongliu9903 /stack_edu_javascripttabular10M<n<100M0 likes114 downloads1y agoHugging Face14Nan-Do /code-search-net-javascript Dataset Card for "code-search-net-javascript" Dataset Summary This dataset is the JavaScript portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the function does. Languages The dataset's comments are in English and the functions are coded in JavaScript Data Splits Train, test, validation labels are… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-javascript.texttext-generation100K<n<1M7 likes109 downloads3y agoHugging Face15h4iku /coconut_javascript2010_preprocessedtext100K<n<1M0 likes102 downloads4y agoHugging Face16CoIR-Retrieval /CodeSearchNet-ccr-javascript-queries-corpus Dataset Card for "CodeSearchNet-ccr-javascript-queries-corpus" More Information needed text100K<n<1M0 likes99 downloads2y agoHugging Face17ars-1 /autotrain-data-javascript-traing-1 AutoTrain Dataset for project: javascript-traing-1 Dataset Description This dataset has been automatically processed by AutoTrain for project javascript-traing-1. Languages The BCP-47 code for the dataset's language is unk. Dataset Structure Data Instances A sample from this dataset looks as follows: [ { "target": "test/NavbarSpec.js", "feat_repo_name": "aabenoja/react-bootstrap", "text": "import React from 'react';\nimport… See the full description on the dataset page: https://huggingface.co/datasets/ars-1/autotrain-data-javascript-traing-1.textsummarization100K<n<1M0 likes96 downloads3y agoHugging Face18AmareshHebbar /leetcode-codegen-javascript LeetCode Code-Gen Dataset — JavaScript 631 rows. Given a problem statement, its input/output examples, and a required algorithm/technique, generate a correct JavaScript solution. Part of a 4-language collection built from the same source: see the sibling Python, Java, C++, and JavaScript datasets. Verification Not execution-verified. There is currently no compiler/runtime harness for this language in the build pipeline (only Python has one). Rows are extracted… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-javascript.texttext-generationn<1K1 likes93 downloads3mo agoHugging Face19CoIR-Retrieval /CodeSearchNet-javascript-qrels Dataset Card for "CodeSearchNet-javascript-qrels" More Information needed text10K<n<100K0 likes89 downloads2y agoHugging Face20CoIR-Retrieval /CodeSearchNet-javascript-queries-corpus Dataset Card for "CodeSearchNet-javascript-queries-corpus" More Information needed text100K<n<1M1 likes75 downloads2y agoHugging Face21nthakur /cornstack-javascript-v1-tevatron-1Mtext100K<n<1M0 likes71 downloads1y agoHugging Face22h4iku /coconut_javascript2010 Dataset Card for CoCoNuT-JavaScript(2010) Dataset Summary Part of the data used to train the models in the "CoCoNuT: Combining Context-Aware Neural Translation Models using Ensemble for Program Repair" paper. These datasets contain raw data extracted from GitHub, GitLab, and Bitbucket, and have neither been shuffled nor tokenized. The year in the dataset’s name is the cutting year that shows the year of the newest commit in the dataset. Languages JavaScript… See the full description on the dataset page: https://huggingface.co/datasets/h4iku/coconut_javascript2010.text1M<n<10M2 likes62 downloads3y agoHugging Face23saurabh5 /rlvr-code-data-JavaScripttext100K<n<1M1 likes61 downloads1y agoHugging Face24semeru /code-text-javascript Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-text/javascript in Semeru CodeXGLUE -- Code-To-Text Task Definition The task is to generate natural language comments for a code, and evaluted by smoothed bleu-4 score. Dataset The dataset we use comes from CodeSearchNet and we filter the dataset as the following:… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-text-javascript.text10K<n<100K10 likes60 downloads4y agoHugging Face25CoIR-Retrieval /CodeSearchNet-ccr-javascript-qrels Dataset Card for "CodeSearchNet-ccr-javascript-qrels" More Information needed text10K<n<100K0 likes59 downloads2y agoHugging Face26finbarr /rlvr-code-data-javascript-editedtext100K<n<1M1 likes43 downloads1y agoHugging Face27Shuu12121 /javascript-treesitter-dedupe-filtered-datasetsV2 Javascript CodeSearch Dataset (Shuu12121/javascript-treesitter-dedupe-filtered-datasetsV2) Dataset Description This dataset contains JavaScript functions and methods paired with their JSDoc comments, extracted from open-source JavaScript repositories on GitHub. It is formatted similarly to the CodeSearchNet challenge dataset. Each entry includes: code: The source code of a javascript function or method. docstring: The docstring or Javadoc associated with the… See the full description on the dataset page: https://huggingface.co/datasets/Shuu12121/javascript-treesitter-dedupe-filtered-datasetsV2.text100K<n<1M0 likes43 downloads1y agoHugging Face28Shuu12121 /github-file-programs-dataset-javascripttext1M<n<10M0 likes35 downloads9mo agoHugging Face29supergoose /buzz_sources_042_javascripttext10K<n<100K0 likes34 downloads2y agoHugging Face30saurabh5 /rlvr-code-data-JavaScript-sfttext100K<n<1M2 likes34 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.