Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SWE-bench /SWE-smith-javatext1K<n<10K0 likes8.1k downloads8mo agoHugging Face02angie-chen55 /javascript-github-codetext10M<n<100M18 likes1.7k downloads4y agoHugging Face03susnato /java_PRstabular100K<n<1M0 likes842 downloads3y agoHugging Face04dmsovetov /codeparrot-javatext10M<n<100M0 likes747 downloads2y agoHugging Face05fyaronskiy /cornstack_java_ru_enThe part of CoRNStack Dataset translated into Russian. Translation was done with Qwen3 model. Samples that satisfy the dual consistency filtering condition (samples where the document_rank is 0 or 1 and document_score > 0.7) were translated. Source code you can find here. For support: fedor.yaronskiy@gmail.com textsentence-similarity1M<n<10M0 likes708 downloads8mo agoHugging Face06tianyang /repobench_java_v1.1 RepoBench v1.1 (Java) Introduction This dataset presents the Java portion of RepoBench v1.1 (ICLR 2024). The data encompasses a collection from GitHub, spanning the period from October 6th to December 31st, 2023. With a commitment to data integrity, we've implemented a deduplication process based on file content against the Stack v2 dataset (coming soon), aiming to mitigate data leakage and memorization concerns. Resources and Links Paper GitHub Dataset… See the full description on the dataset page: https://huggingface.co/datasets/tianyang/repobench_java_v1.1.tabulartext-generation10K<n<100K0 likes548 downloads3y agoHugging Face07LaughingLogits /Stackless_Java_V2 Dataset Summary This is the dataset used for the training of the AP-MAE models, it is a subset of The Heap, we release it for reproducability. tabular1M<n<10M0 likes435 downloads2y agoHugging Face08jitx /Methods2Test_java_unit_test_code Dataset Description Microsoft created this large dataset of Java Junit test cases with its corresponding focal methods. It contains 780k pairs of JUnit test cases and focal methods which were extracted from a total of 91K Java open source project hosted on GitHub. The mapping between test case and focal methods are based heuristics rules and Java developer's best practice. More information could be found here: methods2test Github repo Methods2Test: A dataset of focal methods… See the full description on the dataset page: https://huggingface.co/datasets/jitx/Methods2Test_java_unit_test_code.texttext-generation100K<n<1M18 likes383 downloads3y agoHugging Face09hchautran /javascript-mediumtext100K<n<1M3 likes338 downloads4y agoHugging Face10AlgorithmicResearchGroup /arxiv_java_research_code Dataset Card for "arxiv_java_research_code" More Information needed tabular100K<n<1M1 likes310 downloads3y agoHugging Face11hrishizone /Java-GitHub-Codestext1M<n<10M1 likes267 downloads1y agoHugging Face12MERA-evaluation /JavaTestGen JavaTestGen Task description Java TestGen is a benchmark designed to evaluate code generation models' ability to generate Java unit tests. Tasks involve generating unit tests based on provided Java source code and repository context. Dataset contains 227 tasks. Evaluated skills: Instruction Following, Code Perception, Completion, Testing Contributors: Dmitry Salikhov, Pavel Zadorozhny, Pavel Adamenko, Rodion Levichev, Aidar Valeev, Dmitrii Babaev Motivation… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/JavaTestGen.textn<1K1 likes236 downloads1y agoHugging Face13claudios /java-trace-datasettabular100K<n<1M0 likes230 downloads3y agoHugging Face14JWei05 /SWE-smith-java-6704-filtered-for-problem-statementstext1K<n<10K0 likes229 downloads8mo agoHugging Face15hongliu9903 /stack_edu_javatabular10M<n<100M0 likes225 downloads1y agoHugging Face16hchautran /javascripttext100K<n<1M1 likes221 downloads4y agoHugging Face17ThomasTheMaker /arc-stack-javascripttabular10M<n<100M0 likes214 downloads11mo agoHugging Face18Shuu12121 /github-file-programs-dataset-javatext1M<n<10M0 likes209 downloads9mo agoHugging Face19h4iku /coconut_java2006_preprocessedtext1M<n<10M2 likes206 downloads4y agoHugging Face20microsoft /LCC_java Dataset Card for "LCC_java" More Information needed text100K<n<1M5 likes200 downloads3y agoHugging Face21JoaoJunior /python_java_dataset_APR Dataset Card for "python_java_dataset_APR" More Information needed text1M<n<10M0 likes198 downloads3y agoHugging Face22paulh27 /java_code_api_generationtext1M<n<10M10 likes196 downloads2y agoHugging Face23Nan-Do /code-search-net-java Dataset Card for "code-search-net-java" Dataset Summary This dataset is the Java portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the function does. Languages The dataset's comments are in English and the functions are coded in Java Data Splits Train, test, validation labels are included in the… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-java.textsummarization100K<n<1M4 likes189 downloads3y agoHugging Face24JWei05 /swe_smith_java_qwen3.5_35b_trajs_4369tabular1K<n<10K0 likes173 downloads6mo agoHugging Face25CM /codexglue_code2text_java Dataset Card for "codexglue_code2text_java" More Information needed text100K<n<1M4 likes151 downloads3y agoHugging Face26open-athena /exp_rpt_crosscodeeval-java-qwen3.5-122b-131k-opencode-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/exp_rpt_crosscodeeval-java-qwen3.5-122b-131k-opencode-traces.text1K<n<10K0 likes151 downloads3mo agoHugging Face27JWei05 /SWE-smith-java-6450-filteredtext1K<n<10K0 likes149 downloads8mo agoHugging Face28ammarnasr /the-stack-java-clean Dataset 1: TheStack - Java - Cleaned Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Java, a popular statically typed language. Target Language: Java Dataset Size: Training: 900,000 files Validation: 50,000 files Test: 50,000 files Preprocessing: Selected Java as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-java-clean.tabulartext-generation100K<n<1M13 likes147 downloads3y agoHugging Face29JoaoJunior /java_encoded_processed_APR Dataset Card for "java_encoded_processed_APR" More Information needed text1M<n<10M0 likes140 downloads3y agoHugging Face30hchautran /javascript-smalltext100K<n<1M4 likes139 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.