datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SEC-RPT-Panel
SEC-RPT-Panel v2.0.0
A firm-year research panel for studying related-party-transaction (RPT) fraud
risk, built only from official SEC sources: EDGAR 10-K filings, DEF 14A proxy
statements, XBRL company facts and Accounting and Auditing Enforcement Releases
(AAERs). Published as codewithdark/SEC-RPT-Panel.
Built with sec_rpt_dataset v2.0.0 (git bdc3712) on
2026-10-07.
Tables
File
Content
company_universe.csv
One row per company (CIK): SEC metadata, SIC… See the full description on the dataset page: https://huggingface.co/datasets/codewithdark/SEC-RPT-Panel.tsla-historic-pricesSciCode-Domain-Code
DATA1: Domain-Specific Code Dataset
Dataset Overview
DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code.
Dataset Statistics
Total Datasets: 178 CSV files
Total Data Size: ~115 GB
Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/SciCodePile/SciCode-Domain-Code.harbor-benchus-zip-codes-demographics
Ziplore US ZIP Codes: city, county, coordinates, time zone and Census demographics
Look up any ZIP code's demographics by API instead of loading the file: $5 for 12,000 calls. Buy now → · API key on screen the moment checkout ends · no subscription · 14-day refund if it doesn't work as described · help: cybermax.tools@gmail.com
Need it fresh, filtered or via API? This free file is a snapshot (Census ACS 2020-2024 figures for every US ZIP), last updated 2026-09-24.
Ziplore ZIP… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/us-zip-codes-demographics.test-dataset-v1warn-act-notice-type-codes-crosswalk
WARN Act notice-type codes — the crosswalk
Every US state publishes WARN Act layoff notices with a free-text column saying
what kind of event it is. The statute recognises two: a plant closing and a
mass layoff. Across 48 states that column contains
552 distinct exact strings (531
once you fold case).
This dataset is the crosswalk: one row per raw string, how many notices carry
it, which states emit it, and what it normalizes to.
The finding that matters
520 of… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/warn-act-notice-type-codes-crosswalk.qwen9b-coop-claude-code
qwen9b-coop-claude-code
Two-agent cooperative coding trajectories generated by running
CooperBench in coop mode on
the CooperData task set, using
Qwen/Qwen3.5-9B as the model and Claude Code (claude_code) as the
agent framework. Each pair runs two agents in parallel — one per feature —
coordinating via Redis messaging and a shared git remote.
The matched solo (single-agent) baseline is at
CooperBench/qwen9b-solo-claude-code.
Same task corpus, same model, same agent — only the… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-coop-claude-code.appliancedb-error-codes-repair-database
ApplianceDB: Home Appliance Error Codes & Ranked Repairs
Full dataset: appliancedb.dataengineered.io · $99 one-time (Repair Intelligence Snapshot: commercial licence + SQLite and Parquet builds; the same rows as this sample) → Buy on Stripe · the same sample on Kaggle
Relational database mapping 438 home-appliance error codes across 13 brands and 26 (brand, appliance-type) pairs to 288 ranked repair procedures with DIY difficulty tiers. Every code is identified by its… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/appliancedb-error-codes-repair-database.spotify-top-10k-songsthis list has been extracted from anna's archive : https://annas-archive.li/blog/spotify/spotify-top-10k-songs-table.html
the script used to scrape can be found here : https://gist.github.com/the-code-rider/96838f5d6ff538377776b6ddbb1c633d
Deepseek-V4-Reasoning-Code-2500
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.SciCode-Domain-Code
DATA1: Domain-Specific Code Dataset
Dataset Overview
DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code.
Dataset Statistics
Total Datasets: 178 CSV files
Total Data Size: ~115 GB
Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/SciCode-Domain-Code.crowd-code-dataset-0.1The crowd-code-dataset-0.1 is a raw, unfiltered dataset of fine-grained IDE interactions collected during the development of Jasmine using crowd-code, a VS Code/Cursor extension capturing large parts of the software engineering workflow.
The dataset captures real research engineering workflows (character-level edits, navigation, terminal use, iterative debugging). The crowd-code-dataset-0.1 only includes data from the Jasmine authors. We are actively working on cleaning and curating the full… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/crowd-code-dataset-0.1.rebus-dataset
|🔄 🚍| Re-Bus: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles
Understanding Rebus Puzzles requires a variety of skills such as image recognition, cognitive skills, commonsense reasoning, and multi-step reasoning, making this a challenging task for current Vision-Language Models. In this paper, we present Re-Bus, a large and diverse benchmark of 1,333 English Rebus Puzzles containing different artistic… See the full description on the dataset page: https://huggingface.co/datasets/codergautam/rebus-dataset.qwen9b-solo-claude-code
qwen9b-solo-claude-code
Single-agent coding trajectories generated by running
CooperBench in solo mode on
the CooperData task set, using
Qwen/Qwen3.5-9B as the model and Claude Code (claude_code) as the
agent framework. One agent implements both features in each task.
The matched coop (two-agent) version is at
CooperBench/qwen9b-coop-claude-code.
Same task corpus, same model, same agent — only the coordination differs, so
together they isolate the cooperation deficit.
At a… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-solo-claude-code.neuromoyo-sahara-codeswitch-benchmark
NEUROMOYO — Sahara CodeSwitch Africa Benchmark
🔗 Live Benchmark Results
Interactive benchmark:
https://www.neuromoyo.app/benchmark
This page presents the benchmark results, methodology, model comparisons, robustness analyses, reproducibility information, and limitations for the NEUROMOYO evaluation on African code-switched speech.
🚀 Live NEUROMOYO Demo
Live application:
https://www.neuromoyo.app
The live NEUROMOYO application demonstrates the… See the full description on the dataset page: https://huggingface.co/datasets/prokelly/neuromoyo-sahara-codeswitch-benchmark.goodreads-books
Goodreads Books Dataset
Dataset Description
A comprehensive dataset of books scraped from Goodreads, including ratings, authors, titles, and various book characteristics.
This dataset contains 3045 books with 20 features each, scraped from Goodreads. It's perfect for:
📚 Book recommendation systems
📊 Literary data analysis
🤖 Machine learning projects
📈 Rating prediction models
🔍 Book discovery algorithms
Dataset Structure
Features… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/goodreads-books.qwen9b-coop-claude-code-compressed
qwen9b-coop-claude-code-compressed-ak
Synthetic compressed cooperative agent trajectories derived from
CooperBench/qwen9b-coop-claude-code.
Each raw pair (two LLM coding agents on overlapping features in the same
repo, communicating via Redis messaging + a team git remote) is condensed
into an idealized version: wasted steps dropped, broken submission rituals
fixed, missing cooperation events (coop-send/coop-broadcast/coop-recv,
git fetch/diff team) inserted where the real pair… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen9b-coop-claude-code-compressed.us-vehicle-trouble-codes
US vehicle diagnostic trouble codes, linked to manufacturer bulletins and owner reports
1,142 OBD-II diagnostic trouble codes as they actually appear in US manufacturer service bulletins and NHTSA
owner-complaint filings — with the model years, vehicle components and source documents each code shows up in.
Derived from public NHTSA filings. Regenerated nightly.
DOI: 10.5281/zenodo.22891283 — archived on Zenodo. That is the
concept DOI: it always resolves to the newest version… See the full description on the dataset page: https://huggingface.co/datasets/karsonmadden/us-vehicle-trouble-codes.letterboxd-movies
Letterboxd Movies Dataset
Dataset Description
A comprehensive dataset of movies scraped from Letterboxd, including genres, ratings, runtime, countries, and detailed movie characteristics.
This dataset contains 16246 movies with 28 features each, scraped from Letterboxd. It's perfect for:
🎬 Movie recommendation systems
📊 Film industry analysis
🤖 Machine learning projects
📈 Rating prediction models
🔍 Movie discovery algorithms
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/letterboxd-movies.Saudilang-Code-Switch-Corpus
SCC - Saudilang Code-Switch Corpus
The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "SCC" dataset, which stands for "Saudilang Code-Switch Corpus”.
This dataset contains a transcription of general conversations taken from a YouTube podcast "Thmanyah" that has been transcribed by the National Center for Artificial Intelligence in SDAIA. The data features three episodes covering different domains: investment… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Saudilang-Code-Switch-Corpus.CodeReality
CodeReality: Evaluation Subset - Deliberately Noisy Code Dataset
⚠️ Important Limitations
⚠️ Not Enterprise-Ready: This dataset is deliberately noisy and designed for research only. Contains mixed/unknown licenses, possible secrets, potential security vulnerabilities, duplicate code, and experimental repositories. Requires substantial preprocessing for production use.
Use at your own risk - this is a research dataset for robustness testing and data curation method… See the full description on the dataset page: https://huggingface.co/datasets/vinsblack/CodeReality.AI-Code-Optimization-for-Sustainability-Dataset
AI Code Optimization for Sustainability: Dataset
Refactoring Python Code for Energy-Efficiency using Qwen3: Dataset based on HumanEval, MBPP, and Mercury
📄 Read the Paper | Zenodo Mirror | DOI: 10.5281/zenodo.18377893 | About the author
This dataset is a part of a Master thesis research internship investigating the use of LLMs to optimize Python code for energy efficiency.
The research was conducted as part of the Greenify My Code (GMC) project at the Netherlands Organisation for… See the full description on the dataset page: https://huggingface.co/datasets/BambusControl/AI-Code-Optimization-for-Sustainability-Dataset.fl-local-codes
Florida Local Codes of Ordinances: Publication and Currency, 2026
Where each of Florida's 478 local governments publishes its code of ordinances, who publishes it, and how current that code is. One row per county and per incorporated municipality, covering all 67 counties and all 411 cities, towns and villages. Municipal law is the hardest layer of American law to locate: there is no master index, and each commercial codifier lists only its own clients, so a reader who does not… See the full description on the dataset page: https://huggingface.co/datasets/StepUpLaw/fl-local-codes.gentle-meadow-412d55
gentle-meadow-412d55
Synthetic products test data: 33 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/anthony-code/gentle-meadow-412d55.Deepseek-code
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains… See the full description on the dataset page: https://huggingface.co/datasets/ryen-stuff/Deepseek-code.naics-2022-codes
NAICS 2022 codes
All 2,125 codes in the 2022 U.S. edition of the North American Industry Classification System, from the 20 sectors down to the 1,012 six-digit national industries. Each row has the code, its level, title, parent, sector, and whether the Census Bureau marks it comparable across the U.S., Canada and Mexico.
Two samples come with it:
index_terms_sample.csv has 500 of the 20,373 Census index terms. An index term is a business phrase ("Soybean farming, field and… See the full description on the dataset page: https://huggingface.co/datasets/Graunt/naics-2022-codes.hs-codeonline-retail-refined-datasetSciCode-Domain-Code
DATA1: Domain-Specific Code Dataset
Dataset Overview
DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code.
Dataset Statistics
Total Datasets: 178 CSV files
Total Data Size: ~115 GB
Total Lines… See the full description on the dataset page: https://huggingface.co/datasets/SciCode/SciCode-Domain-Code.
