Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AmazonScience /migration-bench-java-full MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-full.tabulartext-generation1K<n<10K4 likes3.6k downloads1y agoHugging Face02susnato /java_PRstabular100K<n<1M0 likes842 downloads3y agoHugging Face03AmazonScience /migration-bench-java-selected MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-selected.tabulartext-generationn<1K7 likes572 downloads1y agoHugging Face04tianyang /repobench_java_v1.1 RepoBench v1.1 (Java) Introduction This dataset presents the Java portion of RepoBench v1.1 (ICLR 2024). The data encompasses a collection from GitHub, spanning the period from October 6th to December 31st, 2023. With a commitment to data integrity, we've implemented a deduplication process based on file content against the Stack v2 dataset (coming soon), aiming to mitigate data leakage and memorization concerns. Resources and Links Paper GitHub Dataset… See the full description on the dataset page: https://huggingface.co/datasets/tianyang/repobench_java_v1.1.tabulartext-generation10K<n<100K0 likes548 downloads3y agoHugging Face05LaughingLogits /Stackless_Java_V2 Dataset Summary This is the dataset used for the training of the AP-MAE models, it is a subset of The Heap, we release it for reproducability. tabular1M<n<10M0 likes435 downloads2y agoHugging Face06AlgorithmicResearchGroup /arxiv_java_research_code Dataset Card for "arxiv_java_research_code" More Information needed tabular100K<n<1M1 likes310 downloads3y agoHugging Face07LarsEckart /approvaltests-java-sessions Coding agent session traces for LarsEckart/approvaltests-java-sessions This dataset contains redacted coding agent session traces collected while working on git@github.com:approvals/ApprovalTests.Java.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where… See the full description on the dataset page: https://huggingface.co/datasets/LarsEckart/approvaltests-java-sessions.tabulartext-generationn<1K2 likes264 downloads6mo agoHugging Face08iris-sast /CWE-Bench-Java CWE-Bench-Java This repository contains the dataset CWE-Bench-Java presented in the paper LLM-Assisted Static Analysis for Detecting Security Vulnerabilities. At a high level, this dataset contains 120 CVEs spanning 4 CWEs, namely path-traversal, OS-command injection, cross-site scripting, and code-injection. Each CVE includes the buggy and fixed source code of the project, along with the information of the fixed files and functions. We provide the seed information for each CVE in… See the full description on the dataset page: https://huggingface.co/datasets/iris-sast/CWE-Bench-Java.tabular1K<n<10K1 likes257 downloads1y agoHugging Face09claudios /java-trace-datasettabular100K<n<1M0 likes230 downloads3y agoHugging Face10hongliu9903 /stack_edu_javatabular10M<n<100M0 likes225 downloads1y agoHugging Face11ThomasTheMaker /arc-stack-javascripttabular10M<n<100M0 likes214 downloads11mo agoHugging Face12javasoup /koch_test_2025_2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "koch", "total_episodes": 36, "total_frames": 15986, "total_tasks": 1, "total_videos": 72, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:36" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javasoup/koch_test_2025_2.tabularrobotics10K<n<100K0 likes181 downloads1y agoHugging Face13HTJ008 /JavaError-QA JErrRAG-Eval-800 JErrRAG-Eval-800 is the public benchmark release aligned with the paper's final canonical dataset and non-anonymous archival record. This Hugging Face repository contains: java_error_qa_v2/: the canonical public benchmark package paper_online_artifacts/: the paper-facing supplementary artifacts and reproduction bundles SHA256SUMS.txt: release-side hash anchors referenced by the paper Dataset Summary Total records: 800 Split sizes: train=639… See the full description on the dataset page: https://huggingface.co/datasets/HTJ008/JavaError-QA.documentquestion-answeringn<1K2 likes176 downloads3mo agoHugging Face14JWei05 /swe_smith_java_qwen3.5_35b_trajs_4369tabular1K<n<10K0 likes173 downloads6mo agoHugging Face15ammarnasr /the-stack-java-clean Dataset 1: TheStack - Java - Cleaned Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Java, a popular statically typed language. Target Language: Java Dataset Size: Training: 900,000 files Validation: 50,000 files Test: 50,000 files Preprocessing: Selected Java as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-java-clean.tabulartext-generation100K<n<1M13 likes147 downloads3y agoHugging Face16hongliu9903 /stack_edu_javascripttabular10M<n<100M0 likes84 downloads1y agoHugging Face17TheFinAI /github-java-corpusgated github-java-corpus Summary This dataset contains Java source-code text samples prepared for pretraining. Repository TheFinAI/github-java-corpus Required Columns Source: dataset name Date: year Text: the pure text of each sample Token_count: the token count computed with tiktoken Schema Source (string) Date (int32) Text (string) Token_count (int32) Construction The dataset was built from streamed… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/github-java-corpus.tabulartext-generation1M<n<10M0 likes75 downloads3d agoHugging Face18JavaneseHonorifics /Unggah-Ungguh Javanese Honorifics Dataset (Unggah-Ungguh - Released Version) The Javanese language, spoken by over 98 million people, features a distinctive honorific system known as Unggah-Ungguh Basa. In this dataset we present UNGGAH-UNGGUH, a carefully curated dataset designed to encapsulate the nuances of Unggah-Ungguh Basa, the Javanese speech etiquette framework that dictates the choice of words and phrases based on social hierarchy and context. Paper: https://arxiv.org/pdf/2502.20864… See the full description on the dataset page: https://huggingface.co/datasets/JavaneseHonorifics/Unggah-Ungguh.tabular1K<n<10K0 likes67 downloads7mo agoHugging Face19loubnabnl /stack-filtered-pii-1M-java Dataset Card for "stack-filtered-pii-1M-java" More Information needed tabular1M<n<10M0 likes63 downloads4y agoHugging Face20javasoup /act_koch_binky_1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "koch", "total_episodes": 10, "total_frames": 8411, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javasoup/act_koch_binky_1.tabularrobotics1K<n<10K0 likes61 downloads1y agoHugging Face21javadcc /so101_7This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 10, "total_frames": 1901, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javadcc/so101_7.tabularrobotics1K<n<10K0 likes51 downloads11mo agoHugging Face22Zaib /java-vulnerabilitytabular1K<n<10K8 likes44 downloads4y agoHugging Face23javasoup /koch_binky_2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "koch", "total_episodes": 1, "total_frames": 1327, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javasoup/koch_binky_2.tabularrobotics1K<n<10K0 likes43 downloads11mo agoHugging Face24Exqrch /Rebuttal-javanese-pixelgpt Javanese PixelGPT Tokenizer Ablation Dataset Optimized with Font Size 6 and Dynamic Trimming. Tokenizer Schema tok_grapheme: Language-specific Grapheme BPE (izzako/javanese-llama-tokenizer) tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT) tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1) tok_mt5: Google Multilingual Unigram (google/mt5-small) tabular100K<n<1M0 likes43 downloads6mo agoHugging Face25Losa10 /CPT-Java-minecraft-modstabular100K<n<1M0 likes42 downloads2mo agoHugging Face26Reset23 /the-stack-v2-javatabular1M<n<10M0 likes40 downloads2y agoHugging Face27javadcc /so101_1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 10, "total_frames": 2293, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javadcc/so101_1.tabularrobotics1K<n<10K0 likes39 downloads11mo agoHugging Face28athrv /megavul-vulnerability-detection-javatabular10K<n<100K1 likes38 downloads1y agoHugging Face29javasoup /eval_act_koch_test_2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "koch", "total_episodes": 1, "total_frames": 855, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javasoup/eval_act_koch_test_2.tabularrobotics1K<n<10K0 likes37 downloads1y agoHugging Face30javadcc /evorl_screw_147This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "bi_so_follower", "total_episodes": 147, "total_frames": 125473, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:147" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javadcc/evorl_screw_147.tabularrobotics100K<n<1M0 likes37 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.