datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
product-database
Open Food Facts Database
What is 🍊 Open Food Facts?
A food products database
Open Food Facts is a database of food products with ingredients, allergens, nutrition facts and all the tidbits of information we can find on product labels.
Made by everyone
Open Food Facts is a non-profit association of volunteers. 25.000+ contributors like you have added 1.7 million + products from 150 countries using our Android or iPhone app or their… See the full description on the dataset page: https://huggingface.co/datasets/openfoodfacts/product-database.Database
Open Scientific Code Registry (OSCR): the authors' scripts
The code published by the authors of open-access neuroscience papers, as found and
verified by Open Scientific Code Registry (OSCR). Each file is here exactly as it is at the source, at the verified
commit, under the license of its repository.
421,275 unique files (3,782 MB of text) from 8,849 repositories,
in 24 Parquet block(s).
Only files whose repository's license allows redistribution, confirmed by the
repository's… See the full description on the dataset page: https://huggingface.co/datasets/OpenScientificCodeRegistry/Database.ord-data
ord-data
Getting the Data
The datasets live under data/ and are stored with
Git LFS. LFS reads are redirected to the
Hugging Face mirror
via .lfsconfig, so dataset objects are fetched from Hugging
Face's CDN rather than from GitHub's shared (and limited) LFS bandwidth. This is
automatic — you do not need to configure anything.
Option 1: Clone the repository
git clone https://github.com/open-reaction-database/ord-data.git
With Git LFS installed… See the full description on the dataset page: https://huggingface.co/datasets/open-reaction-database/ord-data.jpl-small-body-database
JPL Small-Body Database
Credit: NASA/ESA
Part of a dataset collection on Hugging Face.
Dataset description
Complete catalog of all known asteroids and comets with orbital elements, physical parameters, and discovery metadata. Updated daily from NASA JPL.
The JPL Small-Body Database (SBDB) is the authoritative source for orbital and physical data on all known asteroids, comets, and other small bodies. It is maintained by the Solar System Dynamics group at… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/jpl-small-body-database.crystallography-open-database
Crystallography Open Database (COD) — Full Snapshot
A complete mirror of the Crystallography Open Database (COD) as a single Parquet file, combining all crystallographic metadata with the raw CIF file content in one queryable dataset.
Snapshot Details
Field
Value
Snapshot date
2026-07-06
Metadata fetched
2026-07-06 18:51 (UTC+2) — 533,486 entries
CIF files downloaded
2026-07-06 18:34–21:58 — 533,862 files
Total rows
533,486 (metadata) — 411… See the full description on the dataset page: https://huggingface.co/datasets/LMucko/crystallography-open-database.Telegram-Databasellmops-database
The ZenML LLMOps Database
To learn more about ZenML and our open-source MLOps framework, visit
zenml.io.
Dataset Summary
The LLMOps Database is a comprehensive collection of over 500 real-world
generative AI implementations that showcases how organizations are successfully
deploying Large Language Models (LLMs) in production. The case studies have been
carefully curated to focus on technical depth and practical problem-solving,
with an emphasis on implementation… See the full description on the dataset page: https://huggingface.co/datasets/zenml/llmops-database.Telegram-DatabaseCTIS
Dataset Card for Chinese Traditional Instrument Sound
Original Content
The original dataset is created by [1], with no evaluation provided. The original CTIS dataset contains recordings from 287 varieties of Chinese traditional instruments, reformed Chinese musical instruments, and instruments from ethnic minority groups. Notably, some of these instruments are rarely encountered by the majority of the Chinese populace. The dataset was later utilized by [2] for Chinese… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/CTIS.TORGO-database
The TORGO Database: Acoustic and articulatory speech from speakers with dysarthria
Dataset Summary
This database only includes the short words and restricted sentence portion of the TORGO dataset.
For the full dataset which also includes non-words and unrestricted sentences please see: https://www.cs.toronto.edu/~complingweb/data/TORGO/torgo.html.
Transcripts have been normalized to remove punctuation but casing has been left. Few transcripts only had 'xxx' as text… See the full description on the dataset page: https://huggingface.co/datasets/abnerh/TORGO-database.TrialPanorama-database
Quick start
The easiest way to download the dataset to your local is to use huggingface-cli. The specific command you can use is
huggingface-cli download zifeng-ai/TrialPanorama-database --local-dir LOCAL_DIR --repo-type dataset
where LOCAL_DIR should be replaced with the target directory you want to save your dataset to.
Update history
Aug.4 2025: updated tables with the full set of studies
Dataset website: https://ryanwangzf.github.io/projects/trialpanorama… See the full description on the dataset page: https://huggingface.co/datasets/TrialPanorama/TrialPanorama-database.code-databaseSIWIS_French_Speech_Synthesis_Database
SIWIS French Speech Synthesis Database
This README provides a concise description of the dataset, including its structure, file naming conventions, and known labeling issues. Additionally, suggestions for potential improvements are outlined in the TODO section.
The dataset is distributed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license, permitting its use for any purpose.
For more details about the database design and recording process, please refer… See the full description on the dataset page: https://huggingface.co/datasets/Aviv-anthonnyolime/SIWIS_French_Speech_Synthesis_Database.isbndb-full-database
Dataset Card for "isbndb-annas"
More Information needed
ucs-satellite-database
UCS Satellite Database
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
The Union of Concerned Scientists (UCS) Satellite Database is the most comprehensive publicly available database of operational satellites. Updated roughly quarterly, it includes detailed information about each operational satellite: its name, country of registry, operator, purpose, orbital parameters, launch details, and physical characteristics.
What… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/ucs-satellite-database.database-agent-runs
LibreDB Agent Benchmark
8,199 agent runs · 39 open-weight models served locally, plus one hosted model as a control · 6 task surfaces · 110,711 ledger events · 14,008 refused tool calls
This is the complete measurement record behind the paper What Stops a Small Language Model From
Driving a Database Agent. It is not a scored summary: it is every event the server wrote while the
runs happened, released so that every number in the paper can be recomputed, and disagreed with… See the full description on the dataset page: https://huggingface.co/datasets/libredb/database-agent-runs.database_demoroadmap-databasesUSDA-Phytochemical-Database-JSON
Ethno-API v2.4.0 — Public Sample
Hugging Face hosts a 400-row public sample of Ethno-API v2.4.0: a cleaned and enriched phytochemical data-engineering project derived from the USDA Dr. Duke source data.The full project contains 76,907 records, 2,313 plant species, 24,746 unique chemical entities, and a 16-field public schema with PubMed, ClinicalTrials.gov, ChEMBL, PatentsView, PubChem CID/SMILES, and partner-assisted CID/IUPAC resolution fields.QA-gated public dataset… See the full description on the dataset page: https://huggingface.co/datasets/wirthal1990-tech/USDA-Phytochemical-Database-JSON.gpu-database
GPU Database
Comprehensive GPU specifications database with architecture, manufacturing, API support, performance details, and kernel development specs.
2,824 GPUs across NVIDIA, AMD, and Intel
Part of RightNow — AI-powered code editor for GPU kernel development
Data
Vendor
GPUs
File
NVIDIA
1,286
data/nvidia/all.json
AMD
1,292
data/amd/all.json
Intel
180
data/intel/all.json
All
2,824
data/all-gpus.json
Schema
Each GPU contains up to 55… See the full description on the dataset page: https://huggingface.co/datasets/Jr23xd23/gpu-database.erhu_playing_tech
Dataset Card for Erhu Playing Technique
Original Content
This dataset was created and has been utilized for Erhu playing technique detection by [1], which has not undergone peer review. The original dataset comprises 1,253 Erhu audio clips, all performed by professional Erhu players. These clips were annotated according to three levels, resulting in annotations for four, seven, and 11 categories. Part of the audio data is sourced from the CTIS dataset described earlier.… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/erhu_playing_tech.GZ_IsoTech
Dataset Card for GZ_IsoTech Dataset
Original Content
The dataset is created and used for Guzheng playing technique detection by [1]. The original dataset comprises 2,824 variable-length audio clips showcasing various Guzheng playing techniques. Specifically, 2,328 clips were sourced from virtual sound banks, while 496 clips were performed by a professional Guzheng artist.
The clips are annotated in eight categories, with a Chinese pinyin and Chinese characters written in… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/GZ_IsoTech.slot-database
Slot Machine Database — 5,669 Slots from 58 Providers
Comprehensive dataset of online slot machine metadata covering 58 game providers. Each record includes RTP, volatility, max win multiplier, grid layout, mechanics, themes, features, and bet ranges.
Homepage: slot.report
API: slot.report/api/
Dataset Description
This dataset contains structured metadata for 5,669 online slot machines, making it one of the largest publicly available slot game databases.… See the full description on the dataset page: https://huggingface.co/datasets/slreport/slot-database.gdb_databasesosint-tool-databasefree-global-stock-ticker-database
Free Global Stock Ticker Database
Global stocks and ETFs with listings, identifiers, aliases, and reviewed symbol changes. Maintained by Adanos Software GmbH for ticker detection, identifier resolution, and market-data workflows.
Source and project page: https://adanos.org
GitHub repository: https://github.com/adanos-software/free-ticker-database
Dataset package: adanosorg/free-global-stock-ticker-database
Version: 3.15.0
Contents
File
Rows
Description… See the full description on the dataset page: https://huggingface.co/datasets/adanosorg/free-global-stock-ticker-database.appliancedb-error-codes-repair-database
ApplianceDB: Home Appliance Error Codes & Ranked Repairs
Full dataset: appliancedb.dataengineered.io · $99 one-time (Repair Intelligence Snapshot: commercial licence + SQLite and Parquet builds; the same rows as this sample) → Buy on Stripe · the same sample on Kaggle
Relational database mapping 438 home-appliance error codes across 13 brands and 26 (brand, appliance-type) pairs to 288 ranked repair procedures with DIY difficulty tiers. Every code is identified by its… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/appliancedb-error-codes-repair-database.DatabaseTORGO-database
The TORGO Database: Acoustic and articulatory speech from speakers with dysarthria
Dataset Summary
This database only includes the short words and restricted sentence portion of the TORGO dataset.
For the full dataset which also includes non-words and unrestricted sentences please see: https://www.cs.toronto.edu/~complingweb/data/TORGO/torgo.html.
Transcripts have been normalized to remove punctuation but casing has been left. Few transcripts only had 'xxx' as… See the full description on the dataset page: https://huggingface.co/datasets/SANJAYKISHORE/TORGO-database.fungi_trait_circus_database
fungi_trait_circus_database
大菌輪「Trait Circus」データセット(統制形質)
最終更新日:2025/09/28
重要:データ形式を大幅に更新しました(v2.0)
Languages
Japanese and English
Please do not use this dataset for academic purposes for the time being. (casual use only)
非専門家が作成したデータセットです。学術目的での使用はご遠慮ください。
更新履歴
2025/09/28 (v2.0) - データ構造を全面改訂、Parquet形式に移行、データ量を約2倍に拡充(約400万件)
2025/08/12 (v1.0) - 初回公開版(約180万件)
概要
Atsushi Nakajima(中島淳志)が個人で運営しているWebサイト大菌輪… See the full description on the dataset page: https://huggingface.co/datasets/Atsushi/fungi_trait_circus_database.
