datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
crowd-code-dataset-0.1The crowd-code-dataset-0.1 is a raw, unfiltered dataset of fine-grained IDE interactions collected during the development of Jasmine using crowd-code, a VS Code/Cursor extension capturing large parts of the software engineering workflow.
The dataset captures real research engineering workflows (character-level edits, navigation, terminal use, iterative debugging). The crowd-code-dataset-0.1 only includes data from the Jasmine authors. We are actively working on cleaning and curating the full… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/crowd-code-dataset-0.1.AI-Code-Optimization-for-Sustainability-Dataset
AI Code Optimization for Sustainability: Dataset
Refactoring Python Code for Energy-Efficiency using Qwen3: Dataset based on HumanEval, MBPP, and Mercury
📄 Read the Paper | Zenodo Mirror | DOI: 10.5281/zenodo.18377893 | About the author
This dataset is a part of a Master thesis research internship investigating the use of LLMs to optimize Python code for energy efficiency.
The research was conducted as part of the Greenify My Code (GMC) project at the Netherlands Organisation for… See the full description on the dataset page: https://huggingface.co/datasets/BambusControl/AI-Code-Optimization-for-Sustainability-Dataset.prompt_reverse_engineering_code_dataset_O0_x86_O0code-route-maroc-dataset
🚗 Code de la Route Marocain Dataset (Loi 52-05)
Ce jeu de données regroupe des questions, réponses et textes juridiques formalisés pour l'entraînement de modèles de langage (LLM) sur la réglementation routière au Maroc.
📊 Origine des données
Pipeline NLP Source : Récupéré depuis le projet Kaggle medaymanelkajdouhi/code-route-maroc-nlp.
Format d'export : Fichier export_final.csv converti en data.csv.
🎯 Utilisation
Ce dataset sert de support direct pour le… See the full description on the dataset page: https://huggingface.co/datasets/Zakariae-drabech/code-route-maroc-dataset.Sinhala-English-Code-Mixed-Code-Switched-Hotel-Reviews-Dataset
Sinhala-English Hospitality Review Corpus
This dataset contains 4,133 user reviews collected from Facebook hotel pages across Sri Lanka, annotated at the sentence level for sentiment analysis. The reviews come from 400 hotels across 144 tourist destinations, including Colombo, Kandy, Anuradhapura, Negombo, and other major regions.
Annotation Details
All samples were manually annotated by the author to ensure consistency in sentiment labeling. We welcome volunteers… See the full description on the dataset page: https://huggingface.co/datasets/nlp-dataset/Sinhala-English-Code-Mixed-Code-Switched-Hotel-Reviews-Dataset.prompt_reverse_engineering_code_dataset_O3_arm_O3_advanced_custom_testfinal_recreated_reverse_engineering_code_dataset_O2_x86_O2prompt_reverse_engineering_code_reverse_engineering_code_dataset_O0_x64_O0prompt_reverse_engineering_code_dataset_O2_arm_O2_advanced_custom_testprompt_reverse_engineering_code_dataset_O2_arm_O2_advanced_custom_test_smoketestprompt_reverse_engineering_code_dataset_O0_arm_O0prompt_reverse_engineering_code_dataset_O0_arm_O0_advanced_custom_testprompt_reverse_engineering_code_dataset_O1_arm_O1_advanced_custom_testprompt_reverse_engineering_code_reverse_engineering_code_dataset_O1_mips_O1prompt_reverse_engineering_code_dataset_O2_arm_O2_issueprompt_reverse_engineering_code_reverse_engineering_code_dataset_O3_mips_O3prompt_reverse_engineering_code_dataset_O1_arm_O1prompt_reverse_engineering_code_reverse_engineering_code_dataset_O2_mips_O2prompt_reverse_engineering_code_reverse_engineering_code_dataset_O2_x64_O2prompt_reverse_engineering_code_dataset_O1_x86_O1final_reverse_engineering_code_dataset_O3_x64_O3prompt_reverse_engineering_code_reverse_engineering_code_dataset_O0_mips_O0prompt_rreverse_engineering_code_dataset_O1_x64_O1code_graph_text2cypher_dataset
