Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01codesignal /tsla-historic-pricestabular1K<n<10K2 likes1.3k downloads3y agoHugging Face02SciCodePile /SciCode-Domain-Code DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/SciCodePile/SciCode-Domain-Code.tabulartext-generation1M<n<10M4 likes1.2k downloads7mo agoHugging Face03av-codes /harbor-benchtabular10K<n<100K1 likes602 downloads2y agoHugging Face04CyberMax-tools /us-zip-codes-demographics Ziplore US ZIP Codes: city, county, coordinates, time zone and Census demographics Look up any ZIP code's demographics by API instead of loading the file: $5 for 12,000 calls. Buy now → · API key on screen the moment checkout ends · no subscription · 14-day refund if it doesn't work as described · help: cybermax.tools@gmail.com Need it fresh, filtered or via API? This free file is a snapshot (Census ACS 2020-2024 figures for every US ZIP), last updated 2026-09-24. Ziplore ZIP… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/us-zip-codes-demographics.tabulartabular-regression10K<n<100K0 likes452 downloads9h agoHugging Face05codezerro /test-dataset-v1image100K<n<1M0 likes379 downloads1y agoHugging Face06APProjects /warn-act-notice-type-codes-crosswalk WARN Act notice-type codes — the crosswalk Every US state publishes WARN Act layoff notices with a free-text column saying what kind of event it is. The statute recognises two: a plant closing and a mass layoff. Across 48 states that column contains 552 distinct exact strings (531 once you fold case). This dataset is the crosswalk: one row per raw string, how many notices carry it, which states emit it, and what it normalizes to. The finding that matters 520 of… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/warn-act-notice-type-codes-crosswalk.tabulartabular-classificationn<1K1 likes360 downloads15d agoHugging Face07CooperBench /qwen9b-coop-claude-code qwen9b-coop-claude-code Two-agent cooperative coding trajectories generated by running CooperBench in coop mode on the CooperData task set, using Qwen/Qwen3.5-9B as the model and Claude Code (claude_code) as the agent framework. Each pair runs two agents in parallel — one per feature — coordinating via Redis messaging and a shared git remote. The matched solo (single-agent) baseline is at CooperBench/qwen9b-solo-claude-code. Same task corpus, same model, same agent — only the… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-coop-claude-code.tabulartext-generationn<1K0 likes229 downloads5mo agoHugging Face08Ichlibitiche /appliancedb-error-codes-repair-database ApplianceDB: Home Appliance Error Codes & Ranked Repairs Full dataset: appliancedb.dataengineered.io · $99 one-time (Repair Intelligence Snapshot: commercial licence + SQLite and Parquet builds; the same rows as this sample) → Buy on Stripe · the same sample on Kaggle Relational database mapping 438 home-appliance error codes across 13 brands and 26 (brand, appliance-type) pairs to 288 ranked repair procedures with DIY difficulty tiers. Every code is identified by its… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/appliancedb-error-codes-repair-database.tabular1K<n<10K0 likes186 downloads6d agoHugging Face09code-rider /spotify-top-10k-songsthis list has been extracted from anna's archive : https://annas-archive.li/blog/spotify/spotify-top-10k-songs-table.html the script used to scrape can be found here : https://gist.github.com/the-code-rider/96838f5d6ff538377776b6ddbb1c633d tabular1K<n<10K2 likes180 downloads10mo agoHugging Face10Banaxi-Tech /Deepseek-V4-Reasoning-Code-2500 DeepSeek Reasoning and Code Distillation Dataset This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research. The dataset file is: train.csv It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.tabulartext-generation1K<n<10K13 likes158 downloads5mo agoHugging Face11codergautam /rebus-dataset |🔄 🚍| Re-Bus: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles Understanding Rebus Puzzles requires a variety of skills such as image recognition, cognitive skills, commonsense reasoning, and multi-step reasoning, making this a challenging task for current Vision-Language Models. In this paper, we present Re-Bus, a large and diverse benchmark of 1,333 English Rebus Puzzles containing different artistic… See the full description on the dataset page: https://huggingface.co/datasets/codergautam/rebus-dataset.image1K<n<10K0 likes133 downloads7mo agoHugging Face12liuhangbiao /SciCode-Domain-Code DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/SciCode-Domain-Code.tabulartext-generation1M<n<10M0 likes129 downloads6mo agoHugging Face13CooperBench /qwen9b-solo-claude-code qwen9b-solo-claude-code Single-agent coding trajectories generated by running CooperBench in solo mode on the CooperData task set, using Qwen/Qwen3.5-9B as the model and Claude Code (claude_code) as the agent framework. One agent implements both features in each task. The matched coop (two-agent) version is at CooperBench/qwen9b-coop-claude-code. Same task corpus, same model, same agent — only the coordination differs, so together they isolate the cooperation deficit. At a… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-solo-claude-code.tabulartext-generationn<1K0 likes112 downloads5mo agoHugging Face14PureOne /EVE-SYNRIEL-Witness-Coded-Recursive-Compilation EVE–SYNRIEL Witness-Coded Recursive Compilation for Evidence-Grounded RSI An executable research prototype that compiles action-relevant observations into error-tolerant experiments, retains their evidence ancestry, and applies the same interface to choosing its own task-solving rule. Research v1.0.0 · Hugging Face packaging v1.0.1 · 7 October 2026 Entry point Purpose Manuscript PDF Complete 14-page research report Expert review Proof scope, baseline… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/EVE-SYNRIEL-Witness-Coded-Recursive-Compilation.tabularn<1K0 likes98 downloads3d agoHugging Face15prokelly /neuromoyo-sahara-codeswitch-benchmark NEUROMOYO — Sahara CodeSwitch Africa Benchmark 🔗 Live Benchmark Results Interactive benchmark: https://www.neuromoyo.app/benchmark This page presents the benchmark results, methodology, model comparisons, robustness analyses, reproducibility information, and limitations for the NEUROMOYO evaluation on African code-switched speech. 🚀 Live NEUROMOYO Demo Live application: https://www.neuromoyo.app The live NEUROMOYO application demonstrates the… See the full description on the dataset page: https://huggingface.co/datasets/prokelly/neuromoyo-sahara-codeswitch-benchmark.tabularn<1K0 likes89 downloads19d agoHugging Face16codealchemist01 /goodreads-books Goodreads Books Dataset Dataset Description A comprehensive dataset of books scraped from Goodreads, including ratings, authors, titles, and various book characteristics. This dataset contains 3045 books with 20 features each, scraped from Goodreads. It's perfect for: 📚 Book recommendation systems 📊 Literary data analysis 🤖 Machine learning projects 📈 Rating prediction models 🔍 Book discovery algorithms Dataset Structure Features… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/goodreads-books.tabulartext-classification1K<n<10K0 likes86 downloads1y agoHugging Face17CooperBench /qwen9b-coop-claude-code-compressed qwen9b-coop-claude-code-compressed-ak Synthetic compressed cooperative agent trajectories derived from CooperBench/qwen9b-coop-claude-code. Each raw pair (two LLM coding agents on overlapping features in the same repo, communicating via Redis messaging + a team git remote) is condensed into an idealized version: wasted steps dropped, broken submission rituals fixed, missing cooperation events (coop-send/coop-broadcast/coop-recv, git fetch/diff team) inserted where the real pair… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-coop-claude-code-compressed.tabularn<1K0 likes84 downloads5mo agoHugging Face18p-doom /crowd-code-dataset-0.1The crowd-code-dataset-0.1 is a raw, unfiltered dataset of fine-grained IDE interactions collected during the development of Jasmine using crowd-code, a VS Code/Cursor extension capturing large parts of the software engineering workflow. The dataset captures real research engineering workflows (character-level edits, navigation, terminal use, iterative debugging). The crowd-code-dataset-0.1 only includes data from the Jasmine authors. We are actively working on cleaning and curating the full… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/crowd-code-dataset-0.1.tabular100K<n<1M5 likes80 downloads9mo agoHugging Face19karsonmadden /us-vehicle-trouble-codes US vehicle diagnostic trouble codes, linked to manufacturer bulletins and owner reports 1,142 OBD-II diagnostic trouble codes as they actually appear in US manufacturer service bulletins and NHTSA owner-complaint filings — with the model years, vehicle components and source documents each code shows up in. Derived from public NHTSA filings. Regenerated nightly. DOI: 10.5281/zenodo.22891283 — archived on Zenodo. That is the concept DOI: it always resolves to the newest version… See the full description on the dataset page: https://huggingface.co/datasets/karsonmadden/us-vehicle-trouble-codes.tabulartabular-classification1K<n<10K0 likes73 downloads18d agoHugging Face20SDAIANCAI /Saudilang-Code-Switch-Corpus SCC - Saudilang Code-Switch Corpus The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "SCC" dataset, which stands for "Saudilang Code-Switch Corpus”. This dataset contains a transcription of general conversations taken from a YouTube podcast "Thmanyah" that has been transcribed by the National Center for Artificial Intelligence in SDAIA. The data features three episodes covering different domains: investment… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Saudilang-Code-Switch-Corpus.tabularautomatic-speech-recognition1K<n<10K3 likes68 downloads2y agoHugging Face21codealchemist01 /letterboxd-movies Letterboxd Movies Dataset Dataset Description A comprehensive dataset of movies scraped from Letterboxd, including genres, ratings, runtime, countries, and detailed movie characteristics. This dataset contains 16246 movies with 28 features each, scraped from Letterboxd. It's perfect for: 🎬 Movie recommendation systems 📊 Film industry analysis 🤖 Machine learning projects 📈 Rating prediction models 🔍 Movie discovery algorithms Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/letterboxd-movies.tabulartext-classification10K<n<100K1 likes58 downloads1y agoHugging Face22ronnieaban /hs-codetabulartext-classification1K<n<10K1 likes53 downloads2y agoHugging Face23StepUpLaw /fl-local-codes Florida Local Codes of Ordinances: Publication and Currency, 2026 Where each of Florida's 478 local governments publishes its code of ordinances, who publishes it, and how current that code is. One row per county and per incorporated municipality, covering all 67 counties and all 411 cities, towns and villages. Municipal law is the hardest layer of American law to locate: there is no master index, and each commercial codifier lists only its own clients, so a reader who does not… See the full description on the dataset page: https://huggingface.co/datasets/StepUpLaw/fl-local-codes.tabularn<1K0 likes52 downloads17d agoHugging Face24fai-adh /fon-code-switching-evaluation French-Fon Code-Switching Evaluation Benchmark Overview This dataset is a benchmark designed to evaluate the contextual understanding of small language models in French-Fon code-switching scenarios. The benchmark focuses on situations in which French and Fon (Fongbé) are used within the same interaction, with particular attention to cases where critical information required to answer a question is provided in Fon. The benchmark was developed as part of an academic… See the full description on the dataset page: https://huggingface.co/datasets/fai-adh/fon-code-switching-evaluation.tabularquestion-answeringn<1K0 likes51 downloads2d agoHugging Face25ryen-stuff /Deepseek-code DeepSeek Reasoning and Code Distillation Dataset This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research. The dataset file is: train.csv It contains… See the full description on the dataset page: https://huggingface.co/datasets/ryen-stuff/Deepseek-code.tabulartext-generation1K<n<10K0 likes50 downloads2mo agoHugging Face26anthony-code /gentle-meadow-412d55 gentle-meadow-412d55 Synthetic products test data: 33 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/anthony-code/gentle-meadow-412d55.tabularn<1K0 likes50 downloads1mo agoHugging Face27Graunt /naics-2022-codes NAICS 2022 codes All 2,125 codes in the 2022 U.S. edition of the North American Industry Classification System, from the 20 sectors down to the 1,012 six-digit national industries. Each row has the code, its level, title, parent, sector, and whether the Census Bureau marks it comparable across the U.S., Canada and Mexico. Two samples come with it: index_terms_sample.csv has 500 of the 20,373 Census index terms. An index term is a business phrase ("Soybean farming, field and… See the full description on the dataset page: https://huggingface.co/datasets/Graunt/naics-2022-codes.tabulartext-classification1K<n<10K0 likes50 downloads13d agoHugging Face28vinsblack /CodeReality CodeReality: Evaluation Subset - Deliberately Noisy Code Dataset ⚠️ Important Limitations ⚠️ Not Enterprise-Ready: This dataset is deliberately noisy and designed for research only. Contains mixed/unknown licenses, possible secrets, potential security vulnerabilities, duplicate code, and experimental repositories. Requires substantial preprocessing for production use. Use at your own risk - this is a research dataset for robustness testing and data curation method… See the full description on the dataset page: https://huggingface.co/datasets/vinsblack/CodeReality.tabulartext-generation1K<n<10K1 likes49 downloads1y agoHugging Face29BambusControl /AI-Code-Optimization-for-Sustainability-Dataset AI Code Optimization for Sustainability: Dataset Refactoring Python Code for Energy-Efficiency using Qwen3: Dataset based on HumanEval, MBPP, and Mercury 📄 Read the Paper | Zenodo Mirror | DOI: 10.5281/zenodo.18377893 | About the author This dataset is a part of a Master thesis research internship investigating the use of LLMs to optimize Python code for energy efficiency. The research was conducted as part of the Greenify My Code (GMC) project at the Netherlands Organisation for… See the full description on the dataset page: https://huggingface.co/datasets/BambusControl/AI-Code-Optimization-for-Sustainability-Dataset.tabular10K<n<100K0 likes49 downloads8mo agoHugging Face30codewithdark /online-retail-refined-datasettabular100K<n<1M1 likes39 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.