datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Code_Vulnerability_Labeled_Dataset
Dataset Card for Code_Vulnerability_Labeled_Dataset
Dataset Summary
This dataset provides (code, vulnerability) pairs. The vulnerability field takes values according to the CWE annotation:
CWE
Description
CWE-020
Improper Input Validation
CWE-022
Improper Limitation of a Pathname to a Restricted Directory (“Path Traversal”)
CWE-078
Improper Neutralization of Special Elements used in an OS Command (“OS Command Injection”)
CWE-079
Improper Neutralization of… See the full description on the dataset page: https://huggingface.co/datasets/lemon42-ai/Code_Vulnerability_Labeled_Dataset.crowd-code-dataset-1.0
Install crowd-code 2.0 to help crowd-source the next-generation coding dataset.
crowd-code-dataset-1.0 is an anonymized dataset of fine-grained IDE interactions crowd-sourced across 25 people over the last 6 months using crowd-code 1.0, a VS Code/Cursor extension capturing large parts of the software engineering workflow.
The dataset captures real research engineering workflows (character-level edits, navigation, terminal use, iterative… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/crowd-code-dataset-1.0.Flutter-Code-with-Questions-Dataset-Turkish
Flutter Code with Questions Dataset (Turkish)
📦 Dataset Name: flutter_code_with_questions
Bu veri seti, Flutter framework'ü ile yazılmış kod parçacıkları ve her bir kod parçası için özel olarak üretilmiş detaylı Türkçe soruları içermektedir. Veri seti, kodların eğitim verisi olarak kullanılmasının yanı sıra, LLM (Large Language Model) tabanlı kod anlama ve soru yanıtlama modellerinin geliştirilmesinde kullanılabilir.
📁 Dataset Format
Veri dosyaları CSV… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-Turkish.Flutter-Code-with-Questions-Dataset-English
🧠 Flutter Code with Questions Dataset (English)
This repository contains a high-quality dataset of Flutter-related code snippets paired with automatically generated English technical questions. The dataset is intended for use in training and fine-tuning language models, coding assistants, and educational systems focused on Flutter development.
📂 Dataset Structure
The dataset is divided into 22 CSV files, each containing 200 entries. Every entry includes:
A… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-English.crowd-code-dataset-0.1The crowd-code-dataset-0.1 is a raw, unfiltered dataset of fine-grained IDE interactions collected during the development of Jasmine using crowd-code, a VS Code/Cursor extension capturing large parts of the software engineering workflow.
The dataset captures real research engineering workflows (character-level edits, navigation, terminal use, iterative debugging). The crowd-code-dataset-0.1 only includes data from the Jasmine authors. We are actively working on cleaning and curating the full… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/crowd-code-dataset-0.1.HelpSteer3-general-code-Shift-Qwen-2.5-1.5B-Instruct-dataset-train-generationscrowd-code-dataset-1.0
Install crowd-code 2.0 to help crowd-source the next-generation coding dataset.
crowd-code-dataset-1.0 is an anonymized dataset of fine-grained IDE interactions crowd-sourced across 25 people over the last 6 months using crowd-code 1.0, a VS Code/Cursor extension capturing large parts of the software engineering workflow.
The dataset captures real research engineering workflows (character-level edits, navigation, terminal use, iterative… See the full description on the dataset page: https://huggingface.co/datasets/Mgmgrand420/crowd-code-dataset-1.0.Code-Mixed-Sentiment-Analysis-Dataset
Dataset Generation:
Initially, we select the Amazon Review Dataset as our base data, referenced from Ni et al. (2019)[^1]. We randomly extract 100,000 instances from this dataset. The original labels in this dataset are ratings, scaled from 1 to 5. For our specific task, we categorize them into Positive (rating > 3), Neutral (rating = 3), and Negative (rating < 3), ensuring a balanced number of instances for each label. To generate the synthetic Code-mixed dataset, we apply two… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Sentiment-Analysis-Dataset.AI-Code-Optimization-for-Sustainability-Dataset
AI Code Optimization for Sustainability: Dataset
Refactoring Python Code for Energy-Efficiency using Qwen3: Dataset based on HumanEval, MBPP, and Mercury
📄 Read the Paper | Zenodo Mirror | DOI: 10.5281/zenodo.18377893 | About the author
This dataset is a part of a Master thesis research internship investigating the use of LLMs to optimize Python code for energy efficiency.
The research was conducted as part of the Greenify My Code (GMC) project at the Netherlands Organisation for… See the full description on the dataset page: https://huggingface.co/datasets/BambusControl/AI-Code-Optimization-for-Sustainability-Dataset.Code-Mixed-Offensive-Language-Detection-Dataset
Code-Mixed-Offensive-Language-Identification
This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi.
Dataset Generation:
Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.prompt_reverse_engineering_code_dataset_O0_x86_O0En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset📊 En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset
The En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset is a multilingual dataset of 100,000 product review texts designed for code-mixed sentiment analysis involving English, Bengali, and Roman Bengali.
Each record includes:
🆔 Id
🛒 ProductId
💬 Code-Mixed-Text
💡 Sentiment
The dataset captures diverse linguistic styles, authentic code-mixing, and real-world sentiment patterns from multilingual digital communication.
🌐 Text Distribution
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/DaliaBarua/En-Bn-Code-Mixed-Two-Class-Sentiment-Dataset.code-docstring-datasetcode-route-maroc-dataset
🚗 Code de la Route Marocain Dataset (Loi 52-05)
Ce jeu de données regroupe des questions, réponses et textes juridiques formalisés pour l'entraînement de modèles de langage (LLM) sur la réglementation routière au Maroc.
📊 Origine des données
Pipeline NLP Source : Récupéré depuis le projet Kaggle medaymanelkajdouhi/code-route-maroc-nlp.
Format d'export : Fichier export_final.csv converti en data.csv.
🎯 Utilisation
Ce dataset sert de support direct pour le… See the full description on the dataset page: https://huggingface.co/datasets/Zakariae-drabech/code-route-maroc-dataset.code.datasetprompt_reverse_engineering_code_dataset_O3_arm_O3_advanced_custom_testcode-nomist-llm-datasetCode Nomist 项目微调数据集
用于大模型微调使用,包含格式化后的问题,以及对应答案每个名称使用 | 符号分割。
prompt_reverse_engineering_code_dataset_O2_arm_O2_advanced_custom_testfinal_recreated_reverse_engineering_code_dataset_O2_x86_O2prompt_reverse_engineering_code_reverse_engineering_code_dataset_O1_mips_O1prompt_reverse_engineering_code_reverse_engineering_code_dataset_O0_x64_O0prompt_reverse_engineering_code_dataset_O0_arm_O0prompt_reverse_engineering_code_dataset_O1_arm_O1_advanced_custom_testprompt_reverse_engineering_code_dataset_O2_arm_O2_advanced_custom_test_smoketestEnglish_French_safety_code-mixing_datasetUsing HarmBench promtps as the English baselines
prompt_reverse_engineering_code_dataset_O2_arm_O2_issueprompt_reverse_engineering_code_dataset_O0_arm_O0_advanced_custom_testprompt_reverse_engineering_code_reverse_engineering_code_dataset_O2_mips_O2prompt_reverse_engineering_code_dataset_O1_x86_O1prompt_reverse_engineering_code_dataset_O1_arm_O1
