Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SWE-bench /SWE-smith-javatext1K<n<10K0 likes13k downloads8mo agoHugging Face02AmazonScience /migration-bench-java-full MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-full.tabulartext-generation1K<n<10K4 likes3.9k downloads1y agoHugging Face03XEUIPR /Java-Code-Large-text-onlyJava-Code-Large Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis. By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/XEUIPR/Java-Code-Large-text-only.texttext-generation10M<n<100M0 likes3.3k downloads2d agoHugging Face04bagasshw /n-hypo-java100K<n<1M0 likes3.1k downloads1y agoHugging Face05ajibawa-2023 /JavaScript-Code-LargeJavaScript-Code-Large JavaScript-Code-Large is a large-scale corpus of JavaScript source code comprising around 5 million JavaScript files. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the JavaScript ecosystem. By providing a high-volume, language-specific corpus, JavaScript-Code-Large enables systematic experimentation in JavaScript-focused model training, domain adaptation… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/JavaScript-Code-Large.texttext-generation1M<n<10M34 likes2.4k downloads8mo agoHugging Face06angie-chen55 /javascript-github-codetext10M<n<100M18 likes2.2k downloads4y agoHugging Face07nomic-ai /cornstack-java-v1 CoRNStack Python Dataset The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code, CodeRankEmbed, and CodeRankLLM. CoRNStack Dataset Curation Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered out… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-java-v1.text10M<n<100M3 likes2k downloads2y agoHugging Face08susnato /java_PRstabular100K<n<1M0 likes1k downloads3y agoHugging Face09fyaronskiy /cornstack_java_ru_enThe part of CoRNStack Dataset translated into Russian. Translation was done with Qwen3 model. Samples that satisfy the dual consistency filtering condition (samples where the document_rank is 0 or 1 and document_score > 0.7) were translated. Source code you can find here. For support: fedor.yaronskiy@gmail.com textsentence-similarity1M<n<10M0 likes679 downloads8mo agoHugging Face10LaughingLogits /Stackless_Java_V2 Dataset Summary This is the dataset used for the training of the AP-MAE models, it is a subset of The Heap, we release it for reproducability. tabular1M<n<10M0 likes650 downloads2y agoHugging Face11Kitxuuu /AIXCC-Java-Challenge 🧠 AIXCC Challenge Benchmark – Java Challenges 📌 Overview AIXCC Challenge Benchmark (Java Challenges) is a curated subset of the full AIXCC Challenge Benchmark focused solely on C-based challenges. This benchmark is built upon the official AIXCC Challenge. Each Java challenge consists of either a delta (focused diff) or full (whole project) test case. 📎 Reference Implementation This benchmark is designed to work with our open-source CRS system:… See the full description on the dataset page: https://huggingface.co/datasets/Kitxuuu/AIXCC-Java-Challenge.0 likes625 downloads1y agoHugging Face12javadtaghia /deewaiREALCN-training Repo git@hf.co:datasets/telcom/deewaiREALCN-training DeewaiREALCN Training Data Image–text pairs for training captioning or vision–language models. Each image is a 1024×1024 RGB JPEG portrait with a short English description. Contents data/train/: 9,000 pairs for training. images/: JPEG files (090000.jpg, …). captions.jsonl: one JSON object per line with file_name and text. data/val/: 1,000 pairs for validation with the same layout. Example… See the full description on the dataset page: https://huggingface.co/datasets/javadtaghia/deewaiREALCN-training.image10K<n<100K2 likes619 downloads10mo agoHugging Face13dmsovetov /codeparrot-javatext10M<n<100M0 likes598 downloads2y agoHugging Face14AmazonScience /migration-bench-java-selected MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-selected.tabulartext-generationn<1K7 likes582 downloads1y agoHugging Face15tianyang /repobench_java_v1.1 RepoBench v1.1 (Java) Introduction This dataset presents the Java portion of RepoBench v1.1 (ICLR 2024). The data encompasses a collection from GitHub, spanning the period from October 6th to December 31st, 2023. With a commitment to data integrity, we've implemented a deduplication process based on file content against the Stack v2 dataset (coming soon), aiming to mitigate data leakage and memorization concerns. Resources and Links Paper GitHub Dataset… See the full description on the dataset page: https://huggingface.co/datasets/tianyang/repobench_java_v1.1.tabulartext-generation10K<n<100K0 likes574 downloads3y agoHugging Face16thedruid831 /JavaScript-Code-LargeJavaScript-Code-Large JavaScript-Code-Large is a large-scale corpus of JavaScript source code comprising around 5 million JavaScript files. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the JavaScript ecosystem. By providing a high-volume, language-specific corpus, JavaScript-Code-Large enables systematic experimentation in JavaScript-focused model training, domain adaptation… See the full description on the dataset page: https://huggingface.co/datasets/thedruid831/JavaScript-Code-Large.texttext-generation1M<n<10M0 likes559 downloads8mo agoHugging Face17liuhangbiao /JavaScript-Code-LargeJavaScript-Code-Large JavaScript-Code-Large is a large-scale corpus of JavaScript source code comprising around 5 million JavaScript files. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the JavaScript ecosystem. By providing a high-volume, language-specific corpus, JavaScript-Code-Large enables systematic experimentation in JavaScript-focused model training, domain adaptation… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/JavaScript-Code-Large.texttext-generation1M<n<10M0 likes516 downloads6mo agoHugging Face18jitx /Methods2Test_java_unit_test_code Dataset Description Microsoft created this large dataset of Java Junit test cases with its corresponding focal methods. It contains 780k pairs of JUnit test cases and focal methods which were extracted from a total of 91K Java open source project hosted on GitHub. The mapping between test case and focal methods are based heuristics rules and Java developer's best practice. More information could be found here: methods2test Github repo Methods2Test: A dataset of focal methods… See the full description on the dataset page: https://huggingface.co/datasets/jitx/Methods2Test_java_unit_test_code.texttext-generation100K<n<1M18 likes515 downloads3y agoHugging Face19ajibawa-2023 /Java-Code-LargeJava-Code-Large Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis. By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Java-Code-Large.texttext-generation10M<n<100M34 likes500 downloads8mo agoHugging Face20Ujjwal-Tyagi /JavaScript-Code-LargeJavaScript-Code-Large JavaScript-Code-Large is a large-scale corpus of JavaScript source code comprising around 5 million JavaScript files. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the JavaScript ecosystem. By providing a high-volume, language-specific corpus, JavaScript-Code-Large enables systematic experimentation in JavaScript-focused model training, domain adaptation… See the full description on the dataset page: https://huggingface.co/datasets/Ujjwal-Tyagi/JavaScript-Code-Large.texttext-generation1M<n<10M1 likes469 downloads6mo agoHugging Face21Mo7art /Stack2Graph_KG_java Java StackOverflow Knowledge Graph Summary This Hugging Face dataset repository contains the Java shard of the Stack2Graph StackOverflow Knowledge Graph. Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper. The artifact is optimized for graph-based retrieval, SPARQL analytics, and retrieval-augmented question answering over Stack Overflow content. Stack2Graph source:… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_KG_java.100M<n<1B0 likes433 downloads3mo agoHugging Face22sanjaykz /Java-codetexttext-generation100K<n<1M2 likes411 downloads1y agoHugging Face23JWei05 /SWE-smith-java-6704-filtered-for-problem-statementstext1K<n<10K0 likes409 downloads8mo agoHugging Face24hchautran /javascript-mediumtext100K<n<1M2 likes404 downloads4y agoHugging Face25nomic-ai /cornstack-javascript-v1 CoRNStack Javascript Dataset The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code, CodeRankEmbed, and CodeRankLLM. CoRNStack Dataset Curation Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-javascript-v1.text1M<n<10M5 likes364 downloads2y agoHugging Face26Ujjwal-Tyagi /Java-Code-LargeJava-Code-Large Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis. By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/Ujjwal-Tyagi/Java-Code-Large.texttext-generation10M<n<100M0 likes357 downloads6mo agoHugging Face27jarvis3907 /stark-nano-java-corpus0 likes327 downloads2mo agoHugging Face28MERA-evaluation /JavaTestGen JavaTestGen Task description Java TestGen is a benchmark designed to evaluate code generation models' ability to generate Java unit tests. Tasks involve generating unit tests based on provided Java source code and repository context. Dataset contains 227 tasks. Evaluated skills: Instruction Following, Code Perception, Completion, Testing Contributors: Dmitry Salikhov, Pavel Zadorozhny, Pavel Adamenko, Rodion Levichev, Aidar Valeev, Dmitrii Babaev Motivation… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/JavaTestGen.textn<1K1 likes320 downloads1y agoHugging Face29hrishizone /Java-GitHub-Codestext1M<n<10M1 likes304 downloads1y agoHugging Face30Shuu12121 /github-file-programs-dataset-javatext1M<n<10M0 likes303 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.