Team Ai
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tianyang /repobench_java_v1.1 RepoBench v1.1 (Java) Introduction This dataset presents the Java portion of RepoBench v1.1 (ICLR 2024). The data encompasses a collection from GitHub, spanning the period from October 6th to December 31st, 2023. With a commitment to data integrity, we've implemented a deduplication process based on file content against the Stack v2 dataset (coming soon), aiming to mitigate data leakage and memorization concerns. Resources and Links Paper GitHub Dataset… See the full description on the dataset page: https://huggingface.co/datasets/tianyang/repobench_java_v1.1.tabulartext-generation10K<n<100K0 likes548 downloads3y agoHugging Face02jitx /Methods2Test_java_unit_test_code Dataset Description Microsoft created this large dataset of Java Junit test cases with its corresponding focal methods. It contains 780k pairs of JUnit test cases and focal methods which were extracted from a total of 91K Java open source project hosted on GitHub. The mapping between test case and focal methods are based heuristics rules and Java developer's best practice. More information could be found here: methods2test Github repo Methods2Test: A dataset of focal methods… See the full description on the dataset page: https://huggingface.co/datasets/jitx/Methods2Test_java_unit_test_code.texttext-generation100K<n<1M18 likes383 downloads3y agoHugging Face03Nan-Do /code-search-net-java Dataset Card for "code-search-net-java" Dataset Summary This dataset is the Java portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the function does. Languages The dataset's comments are in English and the functions are coded in Java Data Splits Train, test, validation labels are included in the… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-java.textsummarization100K<n<1M4 likes189 downloads3y agoHugging Face04ammarnasr /the-stack-java-clean Dataset 1: TheStack - Java - Cleaned Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Java, a popular statically typed language. Target Language: Java Dataset Size: Training: 900,000 files Validation: 50,000 files Test: 50,000 files Preprocessing: Selected Java as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-java-clean.tabulartext-generation100K<n<1M13 likes147 downloads3y agoHugging Face05Nan-Do /code-search-net-javascript Dataset Card for "code-search-net-javascript" Dataset Summary This dataset is the JavaScript portion of the CodeSarchNet annotated with a summary column.The code-search-net dataset includes open source functions that include comments found at GitHub.The summary is a short description of what the function does. Languages The dataset's comments are in English and the functions are coded in JavaScript Data Splits Train, test, validation labels are… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/code-search-net-javascript.texttext-generation100K<n<1M7 likes104 downloads3y agoHugging Face06Nan-Do /instructional_code-search-net-javacript Dataset Card for "instructional_code-search-net-javacript" Dataset Summary This is an instructional dataset for JavaScript. The dataset contains two different kind of tasks: Given a piece of code generate a description of what it does. Given a description generate a piece of code that fulfils the description. Languages The dataset is in English. Data Splits There are no splits. Dataset Creation May of 2023 Curation Rationale… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/instructional_code-search-net-javacript.texttext-generation100K<n<1M4 likes101 downloads3y agoHugging Face07AmareshHebbar /leetcode-codegen-javascript LeetCode Code-Gen Dataset — JavaScript 631 rows. Given a problem statement, its input/output examples, and a required algorithm/technique, generate a correct JavaScript solution. Part of a 4-language collection built from the same source: see the sibling Python, Java, C++, and JavaScript datasets. Verification Not execution-verified. There is currently no compiler/runtime harness for this language in the build pipeline (only Python has one). Rows are extracted… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-javascript.texttext-generationn<1K1 likes99 downloads3mo agoHugging Face08AmareshHebbar /leetcode-codegen-java LeetCode Code-Gen Dataset — Java 4068 rows. Given a problem statement, its input/output examples, and a required algorithm/technique, generate a correct Java solution. Part of a 4-language collection built from the same source: see the sibling Python, Java, C++, and JavaScript datasets. Verification Not execution-verified. There is currently no compiler/runtime harness for this language in the build pipeline (only Python has one). Rows are extracted directly from… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/leetcode-codegen-java.texttext-generation1K<n<10K0 likes92 downloads3mo agoHugging Face09Nan-Do /instructional_code-search-net-java Dataset Card for "instructional_code-search-net-java" Dataset Summary This is an instructional dataset for Java. The dataset contains two different kind of tasks: Given a piece of code generate a description of what it does. Given a description generate a piece of code that fulfils the description. Languages The dataset is in English. Data Splits There are no splits. Dataset Creation May of 2023 Curation Rationale This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/instructional_code-search-net-java.texttext-generation100K<n<1M1 likes79 downloads3y agoHugging Face10TheFinAI /github-java-corpusgated github-java-corpus Summary This dataset contains Java source-code text samples prepared for pretraining. Repository TheFinAI/github-java-corpus Required Columns Source: dataset name Date: year Text: the pure text of each sample Token_count: the token count computed with tiktoken Schema Source (string) Date (int32) Text (string) Token_count (int32) Construction The dataset was built from streamed… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/github-java-corpus.tabulartext-generation1M<n<10M0 likes75 downloads3d agoHugging Face11Michael22 /javadoc Java Method to JavaDoc Dataset Overview This dataset is designed for the specific task of fine-tuning a model to generate JavaDoc documentation for Java methods. The dataset contains pairs of Java methods and their corresponding JavaDoc comments, facilitating the model's learning of the relationship between code structure and its descriptive documentation. Data Collection The data is collected from various open-source Java projects hosted on platforms such as… See the full description on the dataset page: https://huggingface.co/datasets/Michael22/javadoc.texttext-generation1M<n<10M3 likes57 downloads2y agoHugging Face12afrizalha /Gatra-1-Javanese GatraOne (Gatra-1) is a synthethic Jawa Krama instruction-tuning dataset, generated by GPT-4. Introducing the Gatra-1 dataset This is a synthetic dataset to fine-tune LLMs into responding in Jawa Krama, the high-register of Javanese language. It is 98% generated using GPT-4, which has very good Jawa Krama capabilities. It is currently a 'beta' version with only 560 input-output prompts. So far, this has been only tested on fine-tuning GPT-3.5 with considerable success.… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/Gatra-1-Javanese.texttext-generationn<1K3 likes50 downloads2y agoHugging Face13bryandts /instruction-dataset-indo-java-sunda-bali-gayo-batak-alas-minang-betawitexttext-generation100K<n<1M0 likes27 downloads2y agoHugging Face14afrizalha /Centhini-1-Javanese Dataset details The dataset comprises 529,575 pretraining examples for both Ngoko and Krama Javanese. The data is almost predominantly translation generated with Deepseek V3. The translations include data from English language Fineweb and a paraphrased translation from Indonesian mc4 dataset. Other examples here include ancient Javanese texts, like Serat Centhini and Babad Tanah Djawi, but also open texts like Javanese wikipedia. To our knowledge, this is the largest easily… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/Centhini-1-Javanese.texttext-generation100K<n<1M3 likes26 downloads2y agoHugging Face15bryandts /alpaca-clean-indo-java-sunda-balitexttext-generation100K<n<1M0 likes23 downloads2y agoHugging Face16afrizalha /Gatra-2-Javanese Dataset details The dataset comprises 36870 prompt-response pairs of Krama Javanese instruction-tuning examples. The data is almost entirely synthetic with minimal human curation. The current dataset supports only single-turn QA, although fine-tuning on instruction-tuned models may allow for transfer of multi-turn capabilities. The prompts are generated by GPT-4o, while the responses are generated by Claude 3 Haiku. The way the data set was generated, the prompt may contain terms in… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/Gatra-2-Javanese.texttext-generation10K<n<100K3 likes19 downloads2y agoHugging Face17Exqrch /javanese-Komodo-pixelgpt Javanese PixelGPT Dataset This dataset contains preprocessed Javanese text data for training PixelGPT models. Dataset Statistics Language: Javanese (jawa) Total samples: 401,542 Train samples: 400,726 Test samples: 816 Tokenizers [TO BE EDITED] Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel representation of aksara text <tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/javanese-Komodo-pixelgpt.tabulartext-generation100K<n<1M0 likes16 downloads8mo agoHugging Face18izzako /javanese-pixelgpt Javanese PixelGPT Dataset This dataset contains preprocessed Javanese text data for training PixelGPT models. Dataset Statistics Language: Javanese (jawa) Total samples: 401,542 Train samples: 400,726 Test samples: 816 Tokenizers Grapheme tokenizer: izzako/javanese-llama-tokenizer LLaMA tokenizer: ernie-research/DualGPT Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel… See the full description on the dataset page: https://huggingface.co/datasets/izzako/javanese-pixelgpt.tabulartext-generation100K<n<1M0 likes12 downloads10mo agoHugging Face19magyar-nlp-szine-java /europarl_hunThe Europarl Hungarian corpus extracted from the proceedings of the European Parliament. (https://www.statmt.org/europarl/) Special HTML entities are removed from the data. We split the raw text files into segments of 2,048 tokens. tokens=17,789,186 words: 12,606,986 sentences: 658,824 texttext-generation10K<n<100K0 likes12 downloads5mo agoHugging Face20rinnieyoung /sea-javanese-cleaned-parquet-v1 SEA Javanese Cleaned Parquet v1 Dataset Summary This dataset is a cleaned Javanese pretraining corpus exported in Hugging Face parquet format. Current public sources used in this release: HuggingFaceFW/fineweb-2 / jav_Latn allenai/c4 / jv afrizalha/Centhini-1-Javanese Cleaning and Deduplication Current pipeline: basic text cleaning short-text filtering repetition filtering rule-based noise filtering document-level exact deduplication across all included… See the full description on the dataset page: https://huggingface.co/datasets/rinnieyoung/sea-javanese-cleaned-parquet-v1.texttext-generation100K<n<1M0 likes11 downloads6mo agoHugging Face21javasop /orbital-schemas Orbital Schemas Dataset Training data for OrbGen - a model that generates valid Orbital schemas (.orb files). Dataset Structure train: 142 examples validation: 16 examples test: 10 examples Features prompt: Natural language description of the desired schema completion: Valid Orbital schema in JSON format domain: Application domain (ecommerce, game, productivity, etc.) complexity: Schema complexity (simple, medium, complex) source: Source of the example… See the full description on the dataset page: https://huggingface.co/datasets/javasop/orbital-schemas.texttext-generationn<1K0 likes7 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.