Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Weyaxi /huggingface-spaces-codes 📊 Dataset Description This dataset comprises code files of Huggingface Spaces that have more than 0 likes as of November 10, 2023. This dataset contains various programming languages totaling in 672 MB of compressed and 2.05 GB of uncompressed data. 📝 Data Fields Field Type Description repository string Huggingface Spaces repository names. sdk string Software Development Kit of the space. license string License type of the space.… See the full description on the dataset page: https://huggingface.co/datasets/Weyaxi/huggingface-spaces-codes.text10K<n<100K12 likes4.8k downloads3y agoHugging Face02codelion /SimpleQA-VerifiedSimpleQA Verified is a 1,000-prompt benchmark for reliably evaluating Large Language Models (LLMs) on short-form factuality and parametric knowledge. The authors from Google DeepMind and Google Research address various limitations of SimpleQA, originally designed by Wei et al. (2024) at OpenAI, including noisy and incorrect labels, topical biases, and question redundancy. SimpleQA Verified was created to provide the research community with a more precise instrument to track genuine progress in… See the full description on the dataset page: https://huggingface.co/datasets/codelion/SimpleQA-Verified.text1K<n<10K4 likes1.9k downloads1y agoHugging Face03codesignal /tsla-historic-pricestabular1K<n<10K2 likes1.5k downloads3y agoHugging Face04SciCodePile /SciCode-Domain-Code DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/SciCodePile/SciCode-Domain-Code.tabulartext-generation1M<n<10M4 likes1.3k downloads7mo agoHugging Face05codesignal /sms-spam-collection SMS Spam Collection v.1 DESCRIPTION The SMS Spam Collection v.1 (hereafter the corpus) is a set of SMS tagged messages that have been collected for SMS Spam research. It contains one set of SMS messages in English of 5,574 messages, tagged acording being ham (legitimate) or spam. 1.1. Compilation This corpus has been collected from free or free for research sources at the Web: A collection of between 425 SMS spam messages extracted manually from the Grumbletext Web… See the full description on the dataset page: https://huggingface.co/datasets/codesignal/sms-spam-collection.text1K<n<10K1 likes1.3k downloads3y agoHugging Face06walk-alone /codeforces-problemstext1K<n<10K3 likes937 downloads2y agoHugging Face07Fraser /dream-coder Program Synthesis Data Generated program synthesis datasets used to train dreamcoder. Currently just supports text & list data. text1K<n<10K6 likes809 downloads4y agoHugging Face08neil-code /dialogsum-test Dataset Card for DIALOGSum Corpus Dataset Description Links Homepage: https://aclanthology.org/2021.findings-acl.449 Repository: https://github.com/cylnlp/dialogsum Paper: https://aclanthology.org/2021.findings-acl.449 Point of Contact: https://huggingface.co/knkarthick Dataset Summary DialogSum is a large-scale dialogue summarization dataset, consisting of 13,460 (Plus 100 holdout data for topic generation) dialogues with corresponding… See the full description on the dataset page: https://huggingface.co/datasets/neil-code/dialogsum-test.textsummarization1K<n<10K15 likes785 downloads3y agoHugging Face09MCES10-Software /Python-Code-Solutions Python Code Solutions Features 1000k of Python Code Solutions for Text Generation and Question Answering Python Coding Problems labelled by topic and difficulty Recommendations Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models. Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic} textquestion-answering10K<n<100K0 likes643 downloads1y agoHugging Face10TigreGotico /arabic-dialects-gold20-code-switch gold20-code-switch Code-switched Arabic sentences with IPA: 20 rows per lect across 33 Arabic lects (the same roster as the sibling TigreGotico/arabic-dialects-gold20). Each row embeds foreign material in a dialectal Arabic frame: inline Latin-script English (and French, for the lects whose live contact language is French), Arabic-script loanwords (سيرفس، كاش، موبايل-class), and Arabizi (Latin-written Arabic with digit gutturals). Columns (TSV, UTF-8, one file per lect): id… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-dialects-gold20-code-switch.texttext-to-speechn<1K0 likes631 downloads3mo agoHugging Face11av-codes /harbor-benchtabular10K<n<100K1 likes578 downloads2y agoHugging Face12CyberMax-tools /us-zip-codes-demographics Ziplore US ZIP Codes: city, county, coordinates, time zone and Census demographics Look up any ZIP code's demographics by API instead of loading the file: $5 for 12,000 calls. Buy now → · API key on screen the moment checkout ends · no subscription · 14-day refund if it doesn't work as described · help: cybermax.tools@gmail.com Need it fresh, filtered or via API? This free file is a snapshot (Census ACS 2020-2024 figures for every US ZIP), last updated 2026-09-24. Ziplore ZIP… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/us-zip-codes-demographics.tabulartabular-regression10K<n<100K0 likes385 downloads3d agoHugging Face13codezerro /test-dataset-v1image100K<n<1M0 likes378 downloads1y agoHugging Face14APProjects /warn-act-notice-type-codes-crosswalk WARN Act notice-type codes — the crosswalk Every US state publishes WARN Act layoff notices with a free-text column saying what kind of event it is. The statute recognises two: a plant closing and a mass layoff. Across 48 states that column contains 552 distinct exact strings (531 once you fold case). This dataset is the crosswalk: one row per raw string, how many notices carry it, which states emit it, and what it normalizes to. The finding that matters 520 of… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/warn-act-notice-type-codes-crosswalk.tabulartabular-classificationn<1K1 likes360 downloads12d agoHugging Face15lemon42-ai /Code_Vulnerability_Labeled_Dataset Dataset Card for Code_Vulnerability_Labeled_Dataset Dataset Summary This dataset provides (code, vulnerability) pairs. The vulnerability field takes values according to the CWE annotation: CWE Description CWE-020 Improper Input Validation CWE-022 Improper Limitation of a Pathname to a Restricted Directory (“Path Traversal”) CWE-078 Improper Neutralization of Special Elements used in an OS Command (“OS Command Injection”) CWE-079 Improper Neutralization of… See the full description on the dataset page: https://huggingface.co/datasets/lemon42-ai/Code_Vulnerability_Labeled_Dataset.texttext-classification1K<n<10K13 likes307 downloads2y agoHugging Face16CodeMixBench /CodeMixBench ℹ️Dataset Card for CodeMixBench [EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation. To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/CodeMixBench/CodeMixBench.texttext-generation10K<n<100K3 likes261 downloads1y agoHugging Face17p-doom /crowd-code-dataset-1.0 Install crowd-code 2.0 to help crowd-source the next-generation coding dataset. crowd-code-dataset-1.0 is an anonymized dataset of fine-grained IDE interactions crowd-sourced across 25 people over the last 6 months using crowd-code 1.0, a VS Code/Cursor extension capturing large parts of the software engineering workflow. The dataset captures real research engineering workflows (character-level edits, navigation, terminal use, iterative… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/crowd-code-dataset-1.0.text1M<n<10M5 likes255 downloads9mo agoHugging Face18huggingface /language_codes_marianMTtextn<1K0 likes238 downloads2y agoHugging Face19NoirZangetsu /Flutter-Code-with-Questions-Dataset-Turkish Flutter Code with Questions Dataset (Turkish) 📦 Dataset Name: flutter_code_with_questions Bu veri seti, Flutter framework'ü ile yazılmış kod parçacıkları ve her bir kod parçası için özel olarak üretilmiş detaylı Türkçe soruları içermektedir. Veri seti, kodların eğitim verisi olarak kullanılmasının yanı sıra, LLM (Large Language Model) tabanlı kod anlama ve soru yanıtlama modellerinin geliştirilmesinde kullanılabilir. 📁 Dataset Format Veri dosyaları CSV… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-Turkish.textquestion-answering1K<n<10K0 likes221 downloads2mo agoHugging Face20NoirZangetsu /Flutter-Code-with-Questions-Dataset-English 🧠 Flutter Code with Questions Dataset (English) This repository contains a high-quality dataset of Flutter-related code snippets paired with automatically generated English technical questions. The dataset is intended for use in training and fine-tuning language models, coding assistants, and educational systems focused on Flutter development. 📂 Dataset Structure The dataset is divided into 22 CSV files, each containing 200 entries. Every entry includes: A… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-English.textquestion-answering1K<n<10K3 likes214 downloads2mo agoHugging Face21code-rider /spotify-top-10k-songsthis list has been extracted from anna's archive : https://annas-archive.li/blog/spotify/spotify-top-10k-songs-table.html the script used to scrape can be found here : https://gist.github.com/the-code-rider/96838f5d6ff538377776b6ddbb1c633d tabular1K<n<10K2 likes205 downloads10mo agoHugging Face22CooperBench /qwen9b-coop-claude-code qwen9b-coop-claude-code Two-agent cooperative coding trajectories generated by running CooperBench in coop mode on the CooperData task set, using Qwen/Qwen3.5-9B as the model and Claude Code (claude_code) as the agent framework. Each pair runs two agents in parallel — one per feature — coordinating via Redis messaging and a shared git remote. The matched solo (single-agent) baseline is at CooperBench/qwen9b-solo-claude-code. Same task corpus, same model, same agent — only the… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-coop-claude-code.tabulartext-generationn<1K0 likes201 downloads4mo agoHugging Face23Ichlibitiche /appliancedb-error-codes-repair-database ApplianceDB: Home Appliance Error Codes & Ranked Repairs Full dataset: appliancedb.dataengineered.io · $99 one-time (Repair Intelligence Snapshot: commercial licence + SQLite and Parquet builds; the same rows as this sample) → Buy on Stripe · the same sample on Kaggle Relational database mapping 438 home-appliance error codes across 13 brands and 26 (brand, appliance-type) pairs to 288 ranked repair procedures with DIY difficulty tiers. Every code is identified by its… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/appliancedb-error-codes-repair-database.tabular1K<n<10K0 likes197 downloads3d agoHugging Face24nisaefendioglu /synthetic-sensitive-data-in-source-code-n300 Synthetic Sensitive Data in Source Code (N=300) Synthetic dataset of 300 source-code / config snippets containing hardcoded secrets and PII.Every sample includes at least one sensitive finding (no clean negatives in the main file). Designed for local masking, secret detection, and OWASP LLM02 — Sensitive Information Disclosure workflows. Version 1.3.6: README Files table documents split extension (split_pattern_custom_n300) and clean CSV. v1.3.4: split files + Dataset Viewer… See the full description on the dataset page: https://huggingface.co/datasets/nisaefendioglu/synthetic-sensitive-data-in-source-code-n300.texttext-classificationn<1K1 likes180 downloads8d agoHugging Face25espejelomar /code_search_net_python_10000_examplestext10K<n<100K14 likes171 downloads5y agoHugging Face26ComplyOnSite /ewc-waste-codes-uk EWC waste code dataset by ComplyOnSite A reusable CSV and JSON copy of the 842 entries in the UK List of Waste, commonly called EWC codes. Each record includes the official code and description, its chapter and sub-chapter, whether the code is hazardous, and the WM3 entry type. CSV — one record per row, suitable for spreadsheets, databases and the Hugging Face table viewer JSON — dataset metadata followed by the same 842 records Method — source lineage, transformations and… See the full description on the dataset page: https://huggingface.co/datasets/ComplyOnSite/ewc-waste-codes-uk.textn<1K0 likes164 downloads3d agoHugging Face27liuhangbiao /SciCode-Domain-Code DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/SciCode-Domain-Code.tabulartext-generation1M<n<10M0 likes163 downloads6mo agoHugging Face28HanxiGuo /CodeMirage CodeMirage Dataset CodeMirage is a comprehensive dataset for evaluating AI-generated code detection capabilities across multiple programming languages and AI models. It contains both human-written code and AI-generated code in normal and paraphrased forms. Dataset Structure The dataset is organized with the following fields: code: The code snippet language: Programming language (Python, Java, JavaScript, CPP, C, CSharp, Go, Ruby, PHP, HTML) source: Source of the code… See the full description on the dataset page: https://huggingface.co/datasets/HanxiGuo/CodeMirage.texttext-classification100K<n<1M4 likes162 downloads1y agoHugging Face29Banaxi-Tech /Deepseek-V4-Reasoning-Code-2500 DeepSeek Reasoning and Code Distillation Dataset This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research. The dataset file is: train.csv It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.tabulartext-generation1K<n<10K13 likes160 downloads5mo agoHugging Face30aadajinkya /python_codes_sampletext10K<n<100K2 likes157 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.