Team Ai
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kenhktsui /code-natural-language-classification-datasetSampling from codeparrot/github-code under more permissive license ['mit', 'apache-2.0', 'bsd-3-clause', 'bsd-2-clause', 'cc0-1.0'] + sampling from minipile. It is intended to be used for training code natural language classifier. texttext-classification1M<n<10M0 likes648 downloads2y agoHugging Face02ObscuraCoder /code-classificationThis dataset collates three class-balanced code classification datasets in the wild, where splits have been stratified by label. The sources are: CodeXGLUE defect detection (sourced from: semeru/code-code-DefectDetection) BigCloneBench clone detection (sourced from: nchen909/bigclonebench-processed) CodeComplex code runtime complexity prediction (sourced from: codeparrot/codecomplex) text1M<n<10M0 likes229 downloads2y agoHugging Face03burtenshaw /PleIAs_common_corpus_code_classificationtext100K<n<1M1 likes94 downloads1y agoHugging Face04kaushik-harsh-99 /Code-Language-Classification Programming Language Classification Dataset A large-scale, balanced dataset for programming language identification from source code snippets. Overview This dataset contains 1.664 million cleaned and labeled source code samples across 16 programming languages, specifically designed for programming language classification and identification tasks. Unlike many code datasets that are primarily built for code generation or retrieval, this dataset was curated… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Code-Language-Classification.texttext-classification1M<n<10M3 likes61 downloads4mo agoHugging Face05NLBSE /nlbse25-code-comment-classificationtabular10K<n<100K0 likes60 downloads2y agoHugging Face06code-switching /topic-classificationtabulartext-classificationn<1K0 likes59 downloads29d agoHugging Face07NLBSE /nlbse27-code-comment-classificationtext100K<n<1M0 likes50 downloads2mo agoHugging Face08Er1111c /Malicious_code_classificationtexttext-classification1K<n<10K5 likes46 downloads2y agoHugging Face09poojaruhal /Code-comment-classification Dataset Card for Code Comment Classification Dataset Summary The dataset contains class comments extracted from various big and diverse open-source projects of three programming languages Java, Smalltalk, and Python. Supported Tasks and Leaderboards Single-label text classification and Multi-label text classification Languages Java, Python, Smalltalk Dataset Structure Data Instances { "class" : "Absy.java", "comment":"*… See the full description on the dataset page: https://huggingface.co/datasets/poojaruhal/Code-comment-classification.text-classification1K<n<10K2 likes32 downloads4y agoHugging Face10LangAGI-Lab /code-dpo-classification Dataset Card for "code-dpo-classification" More Information needed text10K<n<100K3 likes31 downloads3y agoHugging Face11NLBSE /nlbse26-code-comment-classificationtabular1K<n<10K0 likes26 downloads1y agoHugging Face12neil-code /autotrain-data-img-classification AutoTrain Dataset for project: img-classification Dataset Description This dataset has been automatically processed by AutoTrain for project img-classification. Languages The BCP-47 code for the dataset's language is unk. Dataset Structure Data Instances A sample from this dataset looks as follows: [ { "image": "<222x163 RGB PIL image>", "target": 0 }, { "image": "<222x163 RGB PIL image>", "target": 3 }]… See the full description on the dataset page: https://huggingface.co/datasets/neil-code/autotrain-data-img-classification.image-classification0 likes16 downloads3y agoHugging Face13nceyda /YAP470_Code_Classification_Dataset3text10K<n<100K0 likes16 downloads2y agoHugging Face14nceyda /YAP470_Code_Classification_Datasettext10K<n<100K0 likes15 downloads2y agoHugging Face15nceyda /YAP470_Code_Classification_Data2text100K<n<1M0 likes15 downloads2y agoHugging Face16duygutumer /YAP470_Code_Classification_Testtabular1K<n<10K0 likes15 downloads2y agoHugging Face17violetakastreva /line-level-code-vs-text-classificationThis dataset was created for SemEval-2026 Task 13, which focuses on distinguishing machine-generated code from human-written code across multiple programming languages and domains. While the original SemEval task operates at the code snippet level, this dataset provides line-level annotations that enable finer-grained analysis of how code-like and text-like content is distributed within mixed inputs. The dataset is intended to support research in machine-generated code detection, robust… See the full description on the dataset page: https://huggingface.co/datasets/violetakastreva/line-level-code-vs-text-classification.texttext-classification10K<n<100K0 likes11 downloads8mo agoHugging Face18JeswinMS4 /code_text_classification Dataset Card for "code_text_classification" More Information needed textn<1K1 likes10 downloads3y agoHugging Face19nceyda /YAP470_Code_Classification_Dataset2text10K<n<100K0 likes10 downloads2y agoHugging Face20nceyda /YAP470_Code_Classification_Datatext100K<n<1M0 likes9 downloads2y agoHugging Face21nceyda /YAP470_Code_Classification_Final_Datatext10K<n<100K0 likes9 downloads2y agoHugging Face22hfwuxing /Occupational_Classification_Code_of_PRC_2022texttext-classification1K<n<10K0 likes9 downloads1y agoHugging Face23saad2002 /code-review-tone-classificationtext1K<n<10K0 likes8 downloads7mo agoHugging Face24neil-code /autotrain-data-tabular-data-classification0 likes7 downloads3y agoHugging Face25duygutumer /YAP470_Code_Classification_Test_Datatext1K<n<10K0 likes7 downloads2y agoHugging Face26virtualvoidsteve /corrupted_code_classification0 likes4 downloads3y agoHugging Face27manu /code_classification0 likes3 downloads3y agoHugging Face28Halleck45 /code-classificationtext1K<n<10K0 likes2 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.