just-dna-seq/pgs-catalog
PGS Catalog — Scoring Files & Cleaned Metadata Complete mirror of PGS Catalog scoring files converted to Apache Parquet format, together with cleaned and normalised metadata tables. Built automatically by the just-prs pipeline. Last updated: 2026-06-14 16:02 UTC Release Statistics Metric Value Scoring file parquets 5,337 Unique PGS IDs (metadata) 5,337 Genome build GRCh38 Total scoring data size 52.6 GB Release timestamp 2026-06-14T16:02:29Z… See the full description on the dataset page: https://huggingface.co/datasets/just-dna-seq/pgs-catalog.
PGS Catalog — Scoring Files & Cleaned Metadata
Complete mirror of PGS Catalog scoring files converted to Apache Parquet format, together with cleaned and normalised metadata tables.
Built automatically by the `just-prs` pipeline.
Last updated: 2026-06-14 16:02 UTC
Release Statistics
Files
Metadata (data/metadata/)
Cleaned and normalised PGS Catalog metadata tables.
Scoring files (data/scores/)
Individual scoring weight files, one parquet per PGS ID. Each file contains variant-level effect weights with harmonised positions.
Metadata Cleanup Pipeline
The raw PGS Catalog CSVs undergo several transformations:
- Column renaming — verbose PGS column names (e.g.
Polygenic Score (PGS) ID) are mapped to shortsnake_caseequivalents (pgs_id). - Genome build normalization — 9 raw build variants (
hg19,hg37,hg38,NCBI36,hg18,NCBI35,GRCh37,GRCh38,NR) are mapped to canonical values:GRCh37,GRCh38,GRCh36, orNR. - Metric string parsing — performance metrics stored as strings like
"1.55 [1.52,1.58]"or"-0.7 (0.15)"are parsed into structured numeric columns (*_estimate,*_ci_lower,*_ci_upper,*_se) for OR, HR, Beta, AUROC, and C-index. - Performance flattening — performance metrics are joined with evaluation sample sets to include sample size (
n_individuals) and ancestry (ancestry_broad). - Best performance selection —
best_performance.parquetcontains one row per PGS ID, selecting the evaluation with the largest sample size (with a preference for European-ancestry cohorts).
Scoring File Schema
Each scoring parquet contains variant-level data from the PGS Catalog harmonised files:
rsID— dbSNP rsIDchr_name/chr_position— chromosome and position (original)effect_allele/other_allele— alleleseffect_weight— variant weight (beta, log-OR, etc.)hm_chr/hm_pos— harmonised chromosome and position (GRCh38)hm_inferOtherAllele— inferred other allele during harmonisation- PGS header metadata is embedded as file-level Parquet metadata under the key
pgs_catalog_header. - In
data/metadata/scores.parquet,ftp_linkpoints to the parquet mirror in this dataset (data/scores/...parquet) for fast parquet-first loading; original EBI links are preserved inftp_link_ebi.
Usage
Load metadata with polars
import polars as pl
from huggingface_hub import hf_hub_download
path = hf_hub_download(
repo_id="just-dna-seq/pgs-catalog",
filename="data/metadata/scores.parquet",
repo_type="dataset",
)
scores = pl.read_parquet(path)
print(scores.head())Load a scoring file
path = hf_hub_download(
repo_id="just-dna-seq/pgs-catalog",
filename="data/scores/PGS000001_hmPOS_GRCh38.parquet",
repo_type="dataset",
)
weights = pl.read_parquet(path)
print(weights.head())Source & License
- Source: PGS Catalog (EBI / NHGRI)
- License: CC BY 4.0
- Citation: Lambert, S.A. et al. The Polygenic Score Catalog as an open database for reproducibility and systematic evaluation. Nature Genetics 53, 420-425 (2021). doi:10.1038/s41588-021-00783-5
