Team Ai
Datasetpublic

wirthal1990-tech/USDA-Phytochemical-Database-JSON

Ethno-API v2.4.0 — Public Sample Hugging Face hosts a 400-row public sample of Ethno-API v2.4.0: a cleaned and enriched phytochemical data-engineering project derived from the USDA Dr. Duke source data.The full project contains 76,907 records, 2,313 plant species, 24,746 unique chemical entities, and a 16-field public schema with PubMed, ClinicalTrials.gov, ChEMBL, PatentsView, PubChem CID/SMILES, and partner-assisted CID/IUPAC resolution fields.QA-gated public dataset… See the full description on the dataset page: https://huggingface.co/datasets/wirthal1990-tech/USDA-Phytochemical-Database-JSON.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
1likes370downloads
Dataset Card

<p align="center"> <img src="./assets/ethno-api-readme-banner.svg" alt="Ethno-API v2.4.0 — phytochemical data rescue, RAG-ready exports, public dataset metrics, and no medical claims" width="100%"> </p>

Ethno-API v2.4.0 — Public Sample

Hugging Face hosts a 400-row public sample of Ethno-API v2.4.0: a cleaned and enriched phytochemical data-engineering project derived from the USDA Dr. Duke source data. The full project contains 76,907 records, 2,313 plant species, 24,746 unique chemical entities, and a 16-field public schema with PubMed, ClinicalTrials.gov, ChEMBL, PatentsView, PubChem CID/SMILES, and partner-assisted CID/IUPAC resolution fields. QA-gated public dataset sample for data-engineering, retrieval and RAG-readiness workflows. No medical claims. No pharmaceutical validation. No safety or efficacy claims. Research and retrieval use only.

<p align="left"> <a href="https://doi.org/10.5281/zenodo.19660107"><img src="https://zenodo.org/badge/DOI/10.5281/zenodo.19660107.svg" alt="Zenodo DOI"></a> <a href="https://huggingface.co/datasets/wirthal1990-tech/USDA-Phytochemical-Database-JSON"><img src="https://img.shields.io/badge/HuggingFace-Dataset-yellow" alt="Hugging Face Dataset"></a> <a href="https://creativecommons.org/licenses/by/4.0/"><img src="https://img.shields.io/badge/public%20sample-CC%20BY%204.0-lightgrey.svg" alt="Public sample license: CC BY 4.0"></a> </p>


What Is Hosted Here

ItemValue
Hosted fileethno_sample_400.parquet
Hosted splittrain
Public sample size400 rows
Hosted formatParquet
Public sample licenseCC BY 4.0
Full project versionv2.4.0
Full project records76,907
Full project plant species2,313
Full project unique chemical entities24,746
Public schema fields16
Partner-assisted CID/IUPAC records1,197
DOI10.5281/zenodo.19660107

This Hugging Face repository is a public sample and documentation surface, not the full commercial export package.


Load the Public Sample

python
from datasets import load_dataset

ds = load_dataset(
    "wirthal1990-tech/USDA-Phytochemical-Database-JSON",
    split="train",
)

print(ds)
print(ds.column_names)
print(ds[0])

Pandas / Parquet workflow:

python
import pandas as pd

df = pd.read_parquet("hf://datasets/wirthal1990-tech/USDA-Phytochemical-Database-JSON/ethno_sample_400.parquet")

print(df.shape)
print(df[["chemical", "plant_species", "pubchem_cid", "canonical_smiles"]].head())

Dataset Scope

Ethno-API is a machine-readable, enriched version of the USDA Dr. Duke phytochemical and ethnobotanical source data, flattened into a practical JSON/Parquet-oriented structure for:

  • —data rescue and normalization demonstrations
  • —retrieval / RAG ingestion experiments
  • —phytochemical and natural-products data prototypes
  • —QA-gated identifier workflows
  • —data-product portfolio proof for client projects

It is not a medical product, not a clinical validation layer, and not evidence that any compound is safe, effective, therapeutically useful, or suitable for use.


Public Schema — v2.4.0, 16 Fields

FieldTypeCoverageNotes
chemicalstring100%USDA compound label, normalized for flat-table use
plant_speciesstring100%Latin binomial species name
applicationstring / nullpartialSource application / activity context where present
dosagestring / nullpartialSource dosage or concentration text where present; not usage guidance
pubmed_mentions_2026integerenrichment layerPubMed mention-count snapshot
clinical_trials_count_2026integerenrichment layerClinicalTrials.gov study-count snapshot; not clinical validation
chembl_bioactivity_countintegerenrichment layerChEMBL bioactivity measurement count
patent_count_since_2020integer / floatenrichment layerPatentsView / patent-density feature
pubchem_cidinteger / nullapprox. 75–82% depending on export/audit viewPubChem CID from enrichment pipeline
canonical_smilesstring / null57,757 records with SMILESCanonical SMILES retrieved via PubChem
compound_typestring100%Compound classification used for filtering
patent_count_methodstring100%Method label for patent-count derivation
partner_cidinteger / null1,197 recordsPartner-assisted PubChem CID resolution
inchi_keystring / nullsubsetPartner-assisted InChIKey / identifier resolution
iupac_verifiedstring / bool / nullsubsetPartner-assisted identifier-verification state
partner_match_methodstring / nullsubsetMatch method used in partner-resolution file

Partner-resolution file: exports/iupac_cid_resolutions.json — 1,197 partner-assisted CID/IUPAC resolution records.


Enrichment Layers

LayerFunctionLimitation
PubMedMention-count snapshotSearch-density proxy only
ClinicalTrials.govStudy-count snapshotNot clinical validation
ChEMBLBioactivity-count featureCount metadata, not safety or efficacy proof
PatentsViewPatent-density featureInnovation / IP-density proxy
PubChemCID and canonical SMILES enrichmentCoverage depends on successful identifier matching
Partner-assisted resolutionCID/IUPAC identifier resolution subsetIdentifier-level contribution only

QA Pipeline

  1. 1.Normalize USDA source records into a flat analytical schema.
  2. 2.Add external enrichment layers from PubMed, ClinicalTrials.gov, ChEMBL, PatentsView, and PubChem.
  3. 3.Retrieve PubChem CID and canonical SMILES where available.
  4. 4.Add compound classification and patent-count method fields.
  5. 5.Add partner-assisted CID/IUPAC resolution subset.
  6. 6.Run reverse-SMILES QA as a downstream identifier-consistency gate.
  7. 7.Export JSON and Parquet samples for analysis and retrieval experiments.

Reverse-SMILES QA Audit — v2.4.0

VerdictCountInterpretation
validated11,981Strict round-trip pass
plausible8,370Pass with caveats
review_required37,361Visible but not auto-trusted
invalidated45Failed validation; excluded from default retrieval-ready export
insufficient_data19,150No SMILES available
Total76,907Full v2.4.0 input set

Default retrieval-ready rule: exclude invalidated and insufficient_data. Keep review_required visible, but do not auto-trust it.

Export-eligible records by the stated QA rule: 57,712 validated + plausible + review_required = 11,981 + 8,370 + 37,361

Known v2.4.0 QA fixes include thiol-false-alcohol detection, non-carboxylic-acid classifier correction, and strictest-verdict-wins CID tainting logic.


Intended Use

Use caseFit
Retrieval experimentsBuild test corpora for vector search and RAG pipelines
Data cleaning demosShow normalization, enrichment, and QA-gating workflows
Natural-products data prototypesExplore structured phytochemical source data
Portfolio proofDemonstrate data rescue → enrichment → QA → export architecture
Review workflowsInspect identifier-resolution and QA-gate logic on a bounded public sample

Limitations — Read Before Use

  • —Public sample only. This Hub dataset contains a 400-row public sample, not the full 76,907-record export.
  • —No pharmaceutical validation. This dataset does not confirm biological activity, safety, efficacy, dosage relevance, or clinical utility.
  • —No medical claims. Nothing in this repository constitutes medical advice, treatment guidance, product guidance, or usage recommendation.
  • —Source dependency. Baseline relationships reflect USDA/source-database records and may contain historical terminology, sparse annotations, or context limitations.
  • —Coverage gaps. Not all records have PubChem CID, SMILES, InChIKey, dosage/source-concentration text, or activity context.
  • —Partner validation is scoped. The 1,197 partner-assisted records are identifier-resolution work, not full pharmacological validation.
  • —QA verdicts are not clinical verdicts. Reverse-SMILES QA checks identifier consistency, not biological truth.
  • —Snapshot. v2.4.0 reflects a point-in-time enrichment snapshot. External databases can change.

Distribution

ChannelURL
Websitehttps://ethno-api.com
GitHubhttps://github.com/wirthal1990-tech/USDA-Phytochemical-Database-JSON
Hugging Facehttps://huggingface.co/datasets/wirthal1990-tech/USDA-Phytochemical-Database-JSON
Kagglehttps://www.kaggle.com/datasets/alexanderwirth/usda-phytochemical-database-json
Zenodo DOIhttps://doi.org/10.5281/zenodo.19660107

Citation

bibtex
@misc{ethno_api_v24_2026,
  title     = {USDA Phytochemical & Ethnobotanical Database -- Enriched v2.4.0},
  author    = {Wirth, Alexander},
  year      = {2026},
  publisher = {Ethno-API},
  url       = {https://ethno-api.com},
  doi       = {10.5281/zenodo.19660107},
  note      = {76,907 records, 24,746 unique chemical entities, 2,313 plant species}
}

Credits

RoleContributor
Data pipeline and project architectureAlexander Wirth

License

Public sample files in this repository: CC BY 4.0 unless otherwise stated. Full commercial dataset: separate Ethno-API license terms. Code snippets / methodology scripts: MIT only where explicitly marked.