wirthal1990-tech/USDA-Phytochemical-Database-JSON
Ethno-API v2.4.0 — Public Sample Hugging Face hosts a 400-row public sample of Ethno-API v2.4.0: a cleaned and enriched phytochemical data-engineering project derived from the USDA Dr. Duke source data.The full project contains 76,907 records, 2,313 plant species, 24,746 unique chemical entities, and a 16-field public schema with PubMed, ClinicalTrials.gov, ChEMBL, PatentsView, PubChem CID/SMILES, and partner-assisted CID/IUPAC resolution fields.QA-gated public dataset… See the full description on the dataset page: https://huggingface.co/datasets/wirthal1990-tech/USDA-Phytochemical-Database-JSON.
<p align="center"> <img src="./assets/ethno-api-readme-banner.svg" alt="Ethno-API v2.4.0 — phytochemical data rescue, RAG-ready exports, public dataset metrics, and no medical claims" width="100%"> </p>
Ethno-API v2.4.0 — Public Sample
Hugging Face hosts a 400-row public sample of Ethno-API v2.4.0: a cleaned and enriched phytochemical data-engineering project derived from the USDA Dr. Duke source data. The full project contains 76,907 records, 2,313 plant species, 24,746 unique chemical entities, and a 16-field public schema with PubMed, ClinicalTrials.gov, ChEMBL, PatentsView, PubChem CID/SMILES, and partner-assisted CID/IUPAC resolution fields. QA-gated public dataset sample for data-engineering, retrieval and RAG-readiness workflows. No medical claims. No pharmaceutical validation. No safety or efficacy claims. Research and retrieval use only.
<p align="left"> <a href="https://doi.org/10.5281/zenodo.19660107"><img src="https://zenodo.org/badge/DOI/10.5281/zenodo.19660107.svg" alt="Zenodo DOI"></a> <a href="https://huggingface.co/datasets/wirthal1990-tech/USDA-Phytochemical-Database-JSON"><img src="https://img.shields.io/badge/HuggingFace-Dataset-yellow" alt="Hugging Face Dataset"></a> <a href="https://creativecommons.org/licenses/by/4.0/"><img src="https://img.shields.io/badge/public%20sample-CC%20BY%204.0-lightgrey.svg" alt="Public sample license: CC BY 4.0"></a> </p>
What Is Hosted Here
This Hugging Face repository is a public sample and documentation surface, not the full commercial export package.
Load the Public Sample
from datasets import load_dataset
ds = load_dataset(
"wirthal1990-tech/USDA-Phytochemical-Database-JSON",
split="train",
)
print(ds)
print(ds.column_names)
print(ds[0])Pandas / Parquet workflow:
import pandas as pd
df = pd.read_parquet("hf://datasets/wirthal1990-tech/USDA-Phytochemical-Database-JSON/ethno_sample_400.parquet")
print(df.shape)
print(df[["chemical", "plant_species", "pubchem_cid", "canonical_smiles"]].head())Dataset Scope
Ethno-API is a machine-readable, enriched version of the USDA Dr. Duke phytochemical and ethnobotanical source data, flattened into a practical JSON/Parquet-oriented structure for:
- data rescue and normalization demonstrations
- retrieval / RAG ingestion experiments
- phytochemical and natural-products data prototypes
- QA-gated identifier workflows
- data-product portfolio proof for client projects
It is not a medical product, not a clinical validation layer, and not evidence that any compound is safe, effective, therapeutically useful, or suitable for use.
Public Schema — v2.4.0, 16 Fields
Partner-resolution file: exports/iupac_cid_resolutions.json — 1,197 partner-assisted CID/IUPAC resolution records.
Enrichment Layers
QA Pipeline
- Normalize USDA source records into a flat analytical schema.
- Add external enrichment layers from PubMed, ClinicalTrials.gov, ChEMBL, PatentsView, and PubChem.
- Retrieve PubChem CID and canonical SMILES where available.
- Add compound classification and patent-count method fields.
- Add partner-assisted CID/IUPAC resolution subset.
- Run reverse-SMILES QA as a downstream identifier-consistency gate.
- Export JSON and Parquet samples for analysis and retrieval experiments.
Reverse-SMILES QA Audit — v2.4.0
Default retrieval-ready rule: exclude invalidated and insufficient_data. Keep review_required visible, but do not auto-trust it.
Export-eligible records by the stated QA rule: 57,712 validated + plausible + review_required = 11,981 + 8,370 + 37,361
Known v2.4.0 QA fixes include thiol-false-alcohol detection, non-carboxylic-acid classifier correction, and strictest-verdict-wins CID tainting logic.
Intended Use
Limitations — Read Before Use
- Public sample only. This Hub dataset contains a 400-row public sample, not the full 76,907-record export.
- No pharmaceutical validation. This dataset does not confirm biological activity, safety, efficacy, dosage relevance, or clinical utility.
- No medical claims. Nothing in this repository constitutes medical advice, treatment guidance, product guidance, or usage recommendation.
- Source dependency. Baseline relationships reflect USDA/source-database records and may contain historical terminology, sparse annotations, or context limitations.
- Coverage gaps. Not all records have PubChem CID, SMILES, InChIKey, dosage/source-concentration text, or activity context.
- Partner validation is scoped. The 1,197 partner-assisted records are identifier-resolution work, not full pharmacological validation.
- QA verdicts are not clinical verdicts. Reverse-SMILES QA checks identifier consistency, not biological truth.
- Snapshot. v2.4.0 reflects a point-in-time enrichment snapshot. External databases can change.
Distribution
Citation
@misc{ethno_api_v24_2026,
title = {USDA Phytochemical & Ethnobotanical Database -- Enriched v2.4.0},
author = {Wirth, Alexander},
year = {2026},
publisher = {Ethno-API},
url = {https://ethno-api.com},
doi = {10.5281/zenodo.19660107},
note = {76,907 records, 24,746 unique chemical entities, 2,313 plant species}
}Credits
License
Public sample files in this repository: CC BY 4.0 unless otherwise stated. Full commercial dataset: separate Ethno-API license terms. Code snippets / methodology scripts: MIT only where explicitly marked.
