Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01codewithdark /SEC-RPT-Panel SEC-RPT-Panel v2.0.0 A firm-year research panel for studying related-party-transaction (RPT) fraud risk, built only from official SEC sources: EDGAR 10-K filings, DEF 14A proxy statements, XBRL company facts and Accounting and Auditing Enforcement Releases (AAERs). Published as codewithdark/SEC-RPT-Panel. Built with sec_rpt_dataset v2.0.0 (git bdc3712) on 2026-10-07. Tables File Content company_universe.csv One row per company (CIK): SEC metadata, SIC… See the full description on the dataset page: https://huggingface.co/datasets/codewithdark/SEC-RPT-Panel.documenttabular-classification1K<n<10K0 likes3.1k downloads6h agoHugging Face02codesignal /tsla-historic-pricestabular1K<n<10K2 likes1.4k downloads3y agoHugging Face03SciCodePile /SciCode-Domain-Code DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/SciCodePile/SciCode-Domain-Code.tabulartext-generation1M<n<10M4 likes1.3k downloads7mo agoHugging Face04av-codes /harbor-benchtabular10K<n<100K1 likes600 downloads2y agoHugging Face05CyberMax-tools /us-zip-codes-demographics Ziplore US ZIP Codes: city, county, coordinates, time zone and Census demographics Look up any ZIP code's demographics by API instead of loading the file: $5 for 12,000 calls. Buy now → · API key on screen the moment checkout ends · no subscription · 14-day refund if it doesn't work as described · help: cybermax.tools@gmail.com Need it fresh, filtered or via API? This free file is a snapshot (Census ACS 2020-2024 figures for every US ZIP), last updated 2026-09-24. Ziplore ZIP… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/us-zip-codes-demographics.tabulartabular-regression10K<n<100K0 likes389 downloads4d agoHugging Face06codezerro /test-dataset-v1image100K<n<1M0 likes378 downloads1y agoHugging Face07APProjects /warn-act-notice-type-codes-crosswalk WARN Act notice-type codes — the crosswalk Every US state publishes WARN Act layoff notices with a free-text column saying what kind of event it is. The statute recognises two: a plant closing and a mass layoff. Across 48 states that column contains 552 distinct exact strings (531 once you fold case). This dataset is the crosswalk: one row per raw string, how many notices carry it, which states emit it, and what it normalizes to. The finding that matters 520 of… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/warn-act-notice-type-codes-crosswalk.tabulartabular-classificationn<1K1 likes360 downloads13d agoHugging Face08CooperBench /qwen9b-coop-claude-code qwen9b-coop-claude-code Two-agent cooperative coding trajectories generated by running CooperBench in coop mode on the CooperData task set, using Qwen/Qwen3.5-9B as the model and Claude Code (claude_code) as the agent framework. Each pair runs two agents in parallel — one per feature — coordinating via Redis messaging and a shared git remote. The matched solo (single-agent) baseline is at CooperBench/qwen9b-solo-claude-code. Same task corpus, same model, same agent — only the… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-coop-claude-code.tabulartext-generationn<1K0 likes204 downloads4mo agoHugging Face09Ichlibitiche /appliancedb-error-codes-repair-database ApplianceDB: Home Appliance Error Codes & Ranked Repairs Full dataset: appliancedb.dataengineered.io · $99 one-time (Repair Intelligence Snapshot: commercial licence + SQLite and Parquet builds; the same rows as this sample) → Buy on Stripe · the same sample on Kaggle Relational database mapping 438 home-appliance error codes across 13 brands and 26 (brand, appliance-type) pairs to 288 ranked repair procedures with DIY difficulty tiers. Every code is identified by its… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/appliancedb-error-codes-repair-database.tabular1K<n<10K0 likes188 downloads4d agoHugging Face10code-rider /spotify-top-10k-songsthis list has been extracted from anna's archive : https://annas-archive.li/blog/spotify/spotify-top-10k-songs-table.html the script used to scrape can be found here : https://gist.github.com/the-code-rider/96838f5d6ff538377776b6ddbb1c633d tabular1K<n<10K2 likes186 downloads10mo agoHugging Face11Banaxi-Tech /Deepseek-V4-Reasoning-Code-2500 DeepSeek Reasoning and Code Distillation Dataset This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research. The dataset file is: train.csv It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.tabulartext-generation1K<n<10K13 likes159 downloads5mo agoHugging Face12liuhangbiao /SciCode-Domain-Code DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/SciCode-Domain-Code.tabulartext-generation1M<n<10M0 likes145 downloads6mo agoHugging Face13p-doom /crowd-code-dataset-0.1The crowd-code-dataset-0.1 is a raw, unfiltered dataset of fine-grained IDE interactions collected during the development of Jasmine using crowd-code, a VS Code/Cursor extension capturing large parts of the software engineering workflow. The dataset captures real research engineering workflows (character-level edits, navigation, terminal use, iterative debugging). The crowd-code-dataset-0.1 only includes data from the Jasmine authors. We are actively working on cleaning and curating the full… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/crowd-code-dataset-0.1.tabular100K<n<1M5 likes140 downloads9mo agoHugging Face14codergautam /rebus-dataset |🔄 🚍| Re-Bus: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles Understanding Rebus Puzzles requires a variety of skills such as image recognition, cognitive skills, commonsense reasoning, and multi-step reasoning, making this a challenging task for current Vision-Language Models. In this paper, we present Re-Bus, a large and diverse benchmark of 1,333 English Rebus Puzzles containing different artistic… See the full description on the dataset page: https://huggingface.co/datasets/codergautam/rebus-dataset.image1K<n<10K0 likes115 downloads7mo agoHugging Face15CooperBench /qwen9b-solo-claude-code qwen9b-solo-claude-code Single-agent coding trajectories generated by running CooperBench in solo mode on the CooperData task set, using Qwen/Qwen3.5-9B as the model and Claude Code (claude_code) as the agent framework. One agent implements both features in each task. The matched coop (two-agent) version is at CooperBench/qwen9b-coop-claude-code. Same task corpus, same model, same agent — only the coordination differs, so together they isolate the cooperation deficit. At a… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-solo-claude-code.tabulartext-generationn<1K0 likes103 downloads4mo agoHugging Face16prokelly /neuromoyo-sahara-codeswitch-benchmark NEUROMOYO — Sahara CodeSwitch Africa Benchmark 🔗 Live Benchmark Results Interactive benchmark: https://www.neuromoyo.app/benchmark This page presents the benchmark results, methodology, model comparisons, robustness analyses, reproducibility information, and limitations for the NEUROMOYO evaluation on African code-switched speech. 🚀 Live NEUROMOYO Demo Live application: https://www.neuromoyo.app The live NEUROMOYO application demonstrates the… See the full description on the dataset page: https://huggingface.co/datasets/prokelly/neuromoyo-sahara-codeswitch-benchmark.tabularn<1K0 likes89 downloads17d agoHugging Face17codealchemist01 /goodreads-books Goodreads Books Dataset Dataset Description A comprehensive dataset of books scraped from Goodreads, including ratings, authors, titles, and various book characteristics. This dataset contains 3045 books with 20 features each, scraped from Goodreads. It's perfect for: 📚 Book recommendation systems 📊 Literary data analysis 🤖 Machine learning projects 📈 Rating prediction models 🔍 Book discovery algorithms Dataset Structure Features… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/goodreads-books.tabulartext-classification1K<n<10K0 likes87 downloads11mo agoHugging Face18CooperBench /qwen9b-coop-claude-code-compressed qwen9b-coop-claude-code-compressed-ak Synthetic compressed cooperative agent trajectories derived from CooperBench/qwen9b-coop-claude-code. Each raw pair (two LLM coding agents on overlapping features in the same repo, communicating via Redis messaging + a team git remote) is condensed into an idealized version: wasted steps dropped, broken submission rituals fixed, missing cooperation events (coop-send/coop-broadcast/coop-recv, git fetch/diff team) inserted where the real pair… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-coop-claude-code-compressed.tabularn<1K0 likes84 downloads4mo agoHugging Face19karsonmadden /us-vehicle-trouble-codes US vehicle diagnostic trouble codes, linked to manufacturer bulletins and owner reports 1,142 OBD-II diagnostic trouble codes as they actually appear in US manufacturer service bulletins and NHTSA owner-complaint filings — with the model years, vehicle components and source documents each code shows up in. Derived from public NHTSA filings. Regenerated nightly. DOI: 10.5281/zenodo.22891283 — archived on Zenodo. That is the concept DOI: it always resolves to the newest version… See the full description on the dataset page: https://huggingface.co/datasets/karsonmadden/us-vehicle-trouble-codes.tabulartabular-classification1K<n<10K0 likes71 downloads15d agoHugging Face20codealchemist01 /letterboxd-movies Letterboxd Movies Dataset Dataset Description A comprehensive dataset of movies scraped from Letterboxd, including genres, ratings, runtime, countries, and detailed movie characteristics. This dataset contains 16246 movies with 28 features each, scraped from Letterboxd. It's perfect for: 🎬 Movie recommendation systems 📊 Film industry analysis 🤖 Machine learning projects 📈 Rating prediction models 🔍 Movie discovery algorithms Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/letterboxd-movies.tabulartext-classification10K<n<100K1 likes61 downloads11mo agoHugging Face21SDAIANCAI /Saudilang-Code-Switch-Corpus SCC - Saudilang Code-Switch Corpus The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "SCC" dataset, which stands for "Saudilang Code-Switch Corpus”. This dataset contains a transcription of general conversations taken from a YouTube podcast "Thmanyah" that has been transcribed by the National Center for Artificial Intelligence in SDAIA. The data features three episodes covering different domains: investment… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Saudilang-Code-Switch-Corpus.tabularautomatic-speech-recognition1K<n<10K3 likes59 downloads2y agoHugging Face22vinsblack /CodeReality CodeReality: Evaluation Subset - Deliberately Noisy Code Dataset ⚠️ Important Limitations ⚠️ Not Enterprise-Ready: This dataset is deliberately noisy and designed for research only. Contains mixed/unknown licenses, possible secrets, potential security vulnerabilities, duplicate code, and experimental repositories. Requires substantial preprocessing for production use. Use at your own risk - this is a research dataset for robustness testing and data curation method… See the full description on the dataset page: https://huggingface.co/datasets/vinsblack/CodeReality.tabulartext-generation1K<n<10K1 likes53 downloads1y agoHugging Face23BambusControl /AI-Code-Optimization-for-Sustainability-Dataset AI Code Optimization for Sustainability: Dataset Refactoring Python Code for Energy-Efficiency using Qwen3: Dataset based on HumanEval, MBPP, and Mercury 📄 Read the Paper | Zenodo Mirror | DOI: 10.5281/zenodo.18377893 | About the author This dataset is a part of a Master thesis research internship investigating the use of LLMs to optimize Python code for energy efficiency. The research was conducted as part of the Greenify My Code (GMC) project at the Netherlands Organisation for… See the full description on the dataset page: https://huggingface.co/datasets/BambusControl/AI-Code-Optimization-for-Sustainability-Dataset.tabular10K<n<100K0 likes53 downloads8mo agoHugging Face24StepUpLaw /fl-local-codes Florida Local Codes of Ordinances: Publication and Currency, 2026 Where each of Florida's 478 local governments publishes its code of ordinances, who publishes it, and how current that code is. One row per county and per incorporated municipality, covering all 67 counties and all 411 cities, towns and villages. Municipal law is the hardest layer of American law to locate: there is no master index, and each commercial codifier lists only its own clients, so a reader who does not… See the full description on the dataset page: https://huggingface.co/datasets/StepUpLaw/fl-local-codes.tabularn<1K0 likes51 downloads15d agoHugging Face25anthony-code /gentle-meadow-412d55 gentle-meadow-412d55 Synthetic products test data: 33 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/anthony-code/gentle-meadow-412d55.tabularn<1K0 likes50 downloads28d agoHugging Face26ryen-stuff /Deepseek-code DeepSeek Reasoning and Code Distillation Dataset This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research. The dataset file is: train.csv It contains… See the full description on the dataset page: https://huggingface.co/datasets/ryen-stuff/Deepseek-code.tabulartext-generation1K<n<10K0 likes47 downloads2mo agoHugging Face27Graunt /naics-2022-codes NAICS 2022 codes All 2,125 codes in the 2022 U.S. edition of the North American Industry Classification System, from the 20 sectors down to the 1,012 six-digit national industries. Each row has the code, its level, title, parent, sector, and whether the Census Bureau marks it comparable across the U.S., Canada and Mexico. Two samples come with it: index_terms_sample.csv has 500 of the 20,373 Census index terms. An index term is a business phrase ("Soybean farming, field and… See the full description on the dataset page: https://huggingface.co/datasets/Graunt/naics-2022-codes.tabulartext-classification1K<n<10K0 likes47 downloads11d agoHugging Face28ronnieaban /hs-codetabulartext-classification1K<n<10K1 likes46 downloads2y agoHugging Face29codewithdark /online-retail-refined-datasettabular100K<n<1M1 likes41 downloads1y agoHugging Face30SciCode /SciCode-Domain-Codegated DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines… See the full description on the dataset page: https://huggingface.co/datasets/SciCode/SciCode-Domain-Code.tabulartext-generation1M<n<10M0 likes37 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.