Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ysn-rfd /text-dataset-tiny-code-script-py-format USED of tahamajs/medicine_ds_persian for .parquet file USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file USED of Abirate/english_quotes for .jsonl file NEW FILES (05/12/2025) NEW FILES (12/26/2025) NEW FILES (02/15/2026) texttext-generation10K<n<100K3 likes1.5k downloads4mo agoHugging Face02BanglishRev /bangla-english-and-code-mixed-ecommerce-review-dataset BanglishRev: A Large-Scale Bangla-English and Code-mixed Dataset of Product Reviews in E-Commerce Description The BanglishRev dataset is the largest e-commerce product review dataset to date for reviews written in Bengali, English, a mixture of both and Banglish, Bengali words written with English alphabets. The dataset comprises of 1.74 million written reviews from 3.2 million ratings information collected from a total of 128k products being sold in online… See the full description on the dataset page: https://huggingface.co/datasets/BanglishRev/bangla-english-and-code-mixed-ecommerce-review-dataset.image1 likes1.3k downloads2y agoHugging Face03jtatman /python-code-dataset-500k Attention: This dataset is a summary and reformat pulled from github code. You should make your own assumptions based on this. In fact, there is another dataset I formed through parsing that addresses several points: out of 500k python related items, most of them are python-ish, not pythonic the majority of the items here contain excessive licensing inclusion of original code the items here are sometimes not even python but have references There's a whole lot of gpl summaries… See the full description on the dataset page: https://huggingface.co/datasets/jtatman/python-code-dataset-500k.texttext-generation100K<n<1M81 likes895 downloads3y agoHugging Face04kenhktsui /code-natural-language-classification-datasetSampling from codeparrot/github-code under more permissive license ['mit', 'apache-2.0', 'bsd-3-clause', 'bsd-2-clause', 'cc0-1.0'] + sampling from minipile. It is intended to be used for training code natural language classifier. texttext-classification1M<n<10M0 likes648 downloads2y agoHugging Face05aysinghal /code-retrieval-training-datasettext100K<n<1M1 likes366 downloads7mo agoHugging Face06ayshajavd /code-security-vulnerability-dataset Code Security Vulnerability Dataset A curated multi-language dataset of 175,419 code samples labeled with 31 vulnerability classes (30 CWEs + safe) for training multi-label code vulnerability detection models. Labels are mapped to OWASP Top 10 2021 categories. Dataset Details Property Value Total Samples 175,419 Train / Val / Test 140,335 / 17,542 / 17,542 Languages C, C++, Python, JavaScript, Java, PHP, Go Labels 31 (multi-label) Format Parquet with… See the full description on the dataset page: https://huggingface.co/datasets/ayshajavd/code-security-vulnerability-dataset.texttext-classification100K<n<1M8 likes364 downloads6mo agoHugging Face07google-research-datasets /great_codeThe dataset for the variable-misuse task, described in the ICLR 2020 paper 'Global Relational Models of Source Code' [https://openreview.net/forum?id=B1lnbRNtwr] This is the public version of the dataset used in that paper. The original, used to produce the graphs in the paper, could not be open-sourced due to licensing issues. See the public associated code repository [https://github.com/VHellendoorn/ICLR20-Great] for results produced from this dataset. This dataset was generated synthetically from the corpus of Python code in the ETH Py150 Open dataset [https://github.com/google-research-datasets/eth_py150_open].table-to-text1M<n<10M5 likes358 downloads3y agoHugging Face08jugalgajjar /MultiLang-Code-Parser-Dataset MultiLang Code Parser Dataset (MLCPD) MultiLang-Code-Parser-Dataset (MLCPD) provides a large-scale, unified dataset of parsed source code across 10 major programming languages, represented under a universal schema that captures syntax, semantics, and structure in a consistent format. Each entry corresponds to one parsed source file and includes: Language metadata Code-level statistics (lines, errors, AST nodes) Universal Schema JSON (normalized structural representation) MLCPD… See the full description on the dataset page: https://huggingface.co/datasets/jugalgajjar/MultiLang-Code-Parser-Dataset.tabular1M<n<10M2 likes345 downloads1y agoHugging Face09XythicK /code-generation-dataset 📄 Code Generation Dataset A large-scale dataset curated for training and evaluating code generation models. This dataset contains high-quality code snippets, prompts, and metadata suitable for various code synthesis tasks, including prompt completion, function generation, and docstring-to-code translation. 📦 Dataset Summary The code-generation-dataset provides: ✅ Prompts describing coding tasks ✅ Code solutions in Python (or other languages, if applicable) ✅ Metadata… See the full description on the dataset page: https://huggingface.co/datasets/XythicK/code-generation-dataset.10 likes332 downloads1y agoHugging Face10hieunguyenminh /code_contests_dp_datasettabular1K<n<10K0 likes310 downloads2y agoHugging Face11lemon42-ai /Code_Vulnerability_Labeled_Dataset Dataset Card for Code_Vulnerability_Labeled_Dataset Dataset Summary This dataset provides (code, vulnerability) pairs. The vulnerability field takes values according to the CWE annotation: CWE Description CWE-020 Improper Input Validation CWE-022 Improper Limitation of a Pathname to a Restricted Directory (“Path Traversal”) CWE-078 Improper Neutralization of Special Elements used in an OS Command (“OS Command Injection”) CWE-079 Improper Neutralization of… See the full description on the dataset page: https://huggingface.co/datasets/lemon42-ai/Code_Vulnerability_Labeled_Dataset.texttext-classification1K<n<10K13 likes307 downloads2y agoHugging Face12anonymous-submission-dataset-code /TiBuDBimage0 likes264 downloads1mo agoHugging Face13Shuu12121 /owl_code_search_hard_negative_datasets-Pre_kd Owl Code Search Hard Negative Datasets Knowledge Distillation (KD) ベースのハードネガティブ付きコード検索データセットです。コード検索モデルShuu12121/CodeSearch-ModernBERT-Crow-v3-large-len1024-Plusを教師モデルとして、各コメントと説明コメントのペアのデータセットから各クエリに対する関数の類似度スコアを計算し、ハードネガティブ(正解に類似しているが不正解の文書)を付与しています。 概要 目的: コード検索モデルの Contrastive Learning / Knowledge Distillation ファインチューニング 言語: Go, Java, JavaScript, PHP, Python, Ruby, Rust, TypeScript(8言語) 総サンプル数: 4,787,740 データサイズ: 8.73 GB(展開後) / 3.37 GB(ダウンロード時) フォーマット:… See the full description on the dataset page: https://huggingface.co/datasets/Shuu12121/owl_code_search_hard_negative_datasets-Pre_kd.textfeature-extraction10M<n<100M0 likes254 downloads7mo agoHugging Face14p-doom /crowd-code-dataset-1.0 Install crowd-code 2.0 to help crowd-source the next-generation coding dataset. crowd-code-dataset-1.0 is an anonymized dataset of fine-grained IDE interactions crowd-sourced across 25 people over the last 6 months using crowd-code 1.0, a VS Code/Cursor extension capturing large parts of the software engineering workflow. The dataset captures real research engineering workflows (character-level edits, navigation, terminal use, iterative… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/crowd-code-dataset-1.0.text1M<n<10M5 likes245 downloads9mo agoHugging Face15NoirZangetsu /Flutter-Code-with-Questions-Dataset-Turkish Flutter Code with Questions Dataset (Turkish) 📦 Dataset Name: flutter_code_with_questions Bu veri seti, Flutter framework'ü ile yazılmış kod parçacıkları ve her bir kod parçası için özel olarak üretilmiş detaylı Türkçe soruları içermektedir. Veri seti, kodların eğitim verisi olarak kullanılmasının yanı sıra, LLM (Large Language Model) tabanlı kod anlama ve soru yanıtlama modellerinin geliştirilmesinde kullanılabilir. 📁 Dataset Format Veri dosyaları CSV… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-Turkish.textquestion-answering1K<n<10K0 likes214 downloads2mo agoHugging Face16NoirZangetsu /Flutter-Code-with-Questions-Dataset-English 🧠 Flutter Code with Questions Dataset (English) This repository contains a high-quality dataset of Flutter-related code snippets paired with automatically generated English technical questions. The dataset is intended for use in training and fine-tuning language models, coding assistants, and educational systems focused on Flutter development. 📂 Dataset Structure The dataset is divided into 22 CSV files, each containing 200 entries. Every entry includes: A… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-English.textquestion-answering1K<n<10K3 likes214 downloads2mo agoHugging Face17Shuu12121 /owl_code_search_hard_negative_datasets_V2_kdtext10M<n<100M1 likes202 downloads6mo agoHugging Face18NLPC-UOM /Sinhala-English-Code-Mixed-Code-Switched-Dataset Sinhala-English-Code-Mixed-Code-Switched-Dataset This dataset contains 10,000 comments that have been annotated at the sentence level for sentiment analysis, humor detection, hate speech detection, aspect identification, and language identification. The following is the tag scheme. Sentiment - Positive, Negative, Neutral, Conflict Humor - Humorous, Non humorous Hate Speech - Hate-Inducing, Abusive, Not offensive Aspect - Network, Billing or Price, Package, Customer Service, Data… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Sinhala-English-Code-Mixed-Code-Switched-Dataset.text-classification6 likes199 downloads2y agoHugging Face19RnniaSnow /st-code-datasetst_dataset_distillation_by_qwen2.5.jsonl 是由 stgenerate 蒸馏qwen2.514b生成的 /data/complier_improve 中的是在蒸馏的过程中增加编译器在环` st_dataset_local.jsonl 是编译验证通过 st_dpo_dataset.jsonl 是交给ai改正后编译验证通过形成DPO 其他的都是各种原因为通过编译 st_dataset_distillation_by_st_coder_clean 这个适合做反面教材 texttext-generation100K<n<1M0 likes187 downloads7mo agoHugging Face20Omarrran /StackPulse_778K_QnA_Code_dataset 💻 StackOverflow-778K: Multi-Year Developer Q&A Dataset Dataset Summary A large-scale Stack Overflow question dataset containing 778,929 unique questions sampled across 7 years (2015–2022). Each question includes the raw HTML body, plain-text version, tags, score, view count, answer count, and a rich set of derived features for immediate ML use. Collected across 8 sampling runs on Feb 27 2026, deduplicated to 778,929 unique questions with only 2 duplicates removed.… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/StackPulse_778K_QnA_Code_dataset.tabulartext-classification1M<n<10M0 likes182 downloads6mo agoHugging Face21MD-Mushfiqur123 /gianxxxt-code-dataset 🚀 Mushfiqur Giant Code Dataset A unified multi-paradigm coding dataset across 9 premium research pillars. text1M<n<10M0 likes182 downloads1mo agoHugging Face22applied-ai-018 /peacock-data-public-datasets-idc-temp-code0 likes170 downloads2y agoHugging Face23novita /agentic_code_dataset_22Dataset: 22 Real Claude Code Sessions To validate Suffix Decoding's applicability in Agentic Coding scenarios, we collected 22 complete Claude Code session recordings. Dataset Overview Metric Value Collection date December 2025 Total sessions 22 Total conversation turns 17,487 Total runtime 50 hours Total input tokens 6,996,619 Total output tokens 6,094,906 Session Scale Distribution Statistic Min Max Average Conversation turns 273… See the full description on the dataset page: https://huggingface.co/datasets/novita/agentic_code_dataset_22.5 likes163 downloads9mo agoHugging Face24Abirate /code_net_datasettext10K<n<100K3 likes155 downloads5y agoHugging Face25Abirate /code_net_dev_datasettext10K<n<100K1 likes155 downloads5y agoHugging Face26Abirate /code_net_test_final_datasettext10K<n<100K2 likes142 downloads5y agoHugging Face27p-doom /crowd-code-dataset-0.1The crowd-code-dataset-0.1 is a raw, unfiltered dataset of fine-grained IDE interactions collected during the development of Jasmine using crowd-code, a VS Code/Cursor extension capturing large parts of the software engineering workflow. The dataset captures real research engineering workflows (character-level edits, navigation, terminal use, iterative debugging). The crowd-code-dataset-0.1 only includes data from the Jasmine authors. We are actively working on cleaning and curating the full… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/crowd-code-dataset-0.1.tabular100K<n<1M5 likes140 downloads9mo agoHugging Face28sci-datasets /sci-codetext1M<n<10M1 likes134 downloads8mo agoHugging Face29Chemin-AI /advent_of_code_ecv_dataset Advent of Code ECV Dataset Many code generation datasets focus on syntax and structure but lack a strong emphasis on contextual understanding, especially from a storytelling perspective.The Advent of Code ECV (Expanded, Curated, Verified) Dataset addresses this gap by curating and verifying multiple approaches for each challenge from 2024 to provide diverse solutions, comparison of strategies, and better adaptability across different programming paradigms.In addition to training and… See the full description on the dataset page: https://huggingface.co/datasets/Chemin-AI/advent_of_code_ecv_dataset.texttext-generationn<1K7 likes130 downloads2y agoHugging Face30michaelnath /bad_code_to_good_code_dataset Dataset Card for "bad_code_to_good_code_dataset" More Information needed text1M<n<10M0 likes105 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.