Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AmazonScience /migration-bench-java-full MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-full.tabulartext-generation1K<n<10K4 likes3.9k downloads1y agoHugging Face02AmazonScience /migration-bench-java-selected MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-selected.tabulartext-generationn<1K7 likes595 downloads1y agoHugging Face03AmazonScience /migration-bench-java-utg MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-utg.texttext-generation1K<n<10K4 likes223 downloads1y agoHugging Face04iawen /java_unit_testtextn<1K2 likes209 downloads3y agoHugging Face05JavaneseHonorifics /Unggah-Ungguh Javanese Honorifics Dataset (Unggah-Ungguh - Released Version) The Javanese language, spoken by over 98 million people, features a distinctive honorific system known as Unggah-Ungguh Basa. In this dataset we present UNGGAH-UNGGUH, a carefully curated dataset designed to encapsulate the nuances of Unggah-Ungguh Basa, the Javanese speech etiquette framework that dictates the choice of words and phrases based on social hierarchy and context. Paper: https://arxiv.org/pdf/2502.20864… See the full description on the dataset page: https://huggingface.co/datasets/JavaneseHonorifics/Unggah-Ungguh.tabular1K<n<10K0 likes51 downloads7mo agoHugging Face06buelfhood /SOCO_TRAIN_javatextfeature-extraction10K<n<100K0 likes44 downloads1y agoHugging Face07Zaib /java-vulnerabilitytabular1K<n<10K8 likes39 downloads4y agoHugging Face08Mlxa /java_methodstext100K<n<1M0 likes29 downloads3y agoHugging Face09DGurgurov /javanese_sa Sentiment Analysis Data for the Javanese Language Dataset Description: This dataset contains a sentiment analysis data from Wongso et al. (2021). Data Structure: The data was used for the project on injecting external commonsense knowledge into multilingual Large Language Models. Citation: @inproceedings{wongso2021causal, title={Causal and Masked Language Modeling of Javanese Language using Transformer-based Architectures}, author={Wongso, Wilson and Setiawan, David Samuel… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/javanese_sa.texttext-classification10K<n<100K0 likes29 downloads2y agoHugging Face10yuchishi /JavaUnitTest01textn<1K0 likes25 downloads2y agoHugging Face11Speech-data /Javanese-Speech-Dataset 🎧 Javanese Speech Dataset The Javanese Speech Dataset is a structured and scalable speech audio dataset designed to provide high-quality audio data for training modern AI and machine learning models. It includes 85 hours of audio data across 585 files, delivered in MP3 and WAV formats, with a total size of 104 MB. This well-balanced audio dataset offers diverse and representative voice data, with 51% female and 49% male speakers, and an age range spanning from 18 to 50+ years. The… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Javanese-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes24 downloads6mo agoHugging Face12cxycxy000 /migration-bench-java-utg MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/cxycxy000/migration-bench-java-utg.texttext-generation1K<n<10K0 likes22 downloads8mo agoHugging Face13buelfhood /SOCO_javatext10K<n<100K0 likes18 downloads3y agoHugging Face14DGurgurov /javanese_conceptnet ConceptNet Data for the Javanese Language Dataset Description: This dataset contains data extracted from ConceptNet using the dedicated module for fetching knowledge from the graph, available on GitHub. Data Structure: The data is converted from triplets into natural text using a pre-defined relationship mapping and split into training and validation sets. It was used for training language adapters for the project aimed at injecting external commonsense knowledge into multilingual… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/javanese_conceptnet.text1K<n<10K1 likes18 downloads3y agoHugging Face15rauf8888 /Synth_Java-Interviewtext1K<n<10K0 likes18 downloads2y agoHugging Face16yukebrillianth /east_java_dialect_instruct Complaints From The East Javanese Dialect community This dataset created manually by humans with reference to public complaints in the comments column of the local government's Instagram account and another platform like X and TikTok Comments. textquestion-answeringn<1K0 likes15 downloads1y agoHugging Face17cmaeti /genesys_api_javascript_alpacatext1K<n<10K0 likes14 downloads2y agoHugging Face18rufimelo /gbug-javaFrom https://github.com/ASSERT-KTH/repairllama/tree/main/benchmarks textn<1K0 likes13 downloads2y agoHugging Face19Satya25 /cobol-to-java-datasettextn<1K0 likes9 downloads3y agoHugging Face20Iftitahu /javanese_instruct_storiesgatedA dataset of parallel translation-based instructions for Javanese language as a target language. Materials are taken from randomly selected children stories at https://storyweaver.org.in, under CC-By-SA-4.0 license. The template IDs are: (1, 'Terjemahno penggalan teks crito ing ngisor iki saka Bahasa Inggris dadi teks crito ing Basa Jawa:', 'Terjemahane utawa padanan teks crito kasebut ing Basa Jawa yaiku:'), (2, 'Terjemahno penggalan teks crito ing ngisor iki saka Bahasa Indonesia dadi teks… See the full description on the dataset page: https://huggingface.co/datasets/Iftitahu/javanese_instruct_stories.texttranslationn<1K2 likes8 downloads3y agoHugging Face21ShadowSnow /java-testtextn<1K0 likes7 downloads3y agoHugging Face22sidharthsingh1892 /cobol_to_java_new_datasettextn<1K0 likes6 downloads3y agoHugging Face23razat-ag /embold_hf_javatextn<1K0 likes6 downloads3y agoHugging Face24javadr /common_voice_16_0_fa_pseudo_labelledtext10K<n<100K0 likes6 downloads3y agoHugging Face25cmaeti /genesys_api_java_alpacatext1K<n<10K0 likes6 downloads2y agoHugging Face26javalove93 /sentiment-analysis-datasettextn<1K1 likes6 downloads2y agoHugging Face27LucianoPrd1 /java-code-minitextn<1K0 likes6 downloads2y agoHugging Face28ShastriPranav /Java_QBtextn<1K0 likes5 downloads3y agoHugging Face29qqqqq1 /Javaragastextn<1K0 likes5 downloads3y agoHugging Face30rmdhirr /javanese-1k-dpo-cleanedtabular1K<n<10K0 likes5 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.