Team Ai
Datasetpublic

just-dna-seq/pgs-catalog

PGS Catalog — Scoring Files & Cleaned Metadata Complete mirror of PGS Catalog scoring files converted to Apache Parquet format, together with cleaned and normalised metadata tables. Built automatically by the just-prs pipeline. Last updated: 2026-06-14 16:02 UTC Release Statistics Metric Value Scoring file parquets 5,337 Unique PGS IDs (metadata) 5,337 Genome build GRCh38 Total scoring data size 52.6 GB Release timestamp 2026-06-14T16:02:29Z… See the full description on the dataset page: https://huggingface.co/datasets/just-dna-seq/pgs-catalog.

sourceHugging Facecc-by-4.0updated 13d agoView on Hugging Face
0likes184downloads
Dataset Card

PGS Catalog — Scoring Files & Cleaned Metadata

Complete mirror of PGS Catalog scoring files converted to Apache Parquet format, together with cleaned and normalised metadata tables.

Built automatically by the `just-prs` pipeline.

Last updated: 2026-06-14 16:02 UTC

Release Statistics

MetricValue
Scoring file parquets5,337
Unique PGS IDs (metadata)5,337
Genome buildGRCh38
Total scoring data size52.6 GB
Release timestamp2026-06-14T16:02:29Z

Files

Metadata (data/metadata/)

Cleaned and normalised PGS Catalog metadata tables.

FileRowsColumnsKey columns
metadata/scores.parquet5,33715pgs_id, name, trait_reported, trait_efo, trait_efo_id, genome_build, n_variants, weight_type, ... (15 total)
metadata/performance.parquet21,19632ppm_id, pgs_id, pss_id, pgp_id, trait_reported, covariates, pmid, doi, ... (32 total)
metadata/best_performance.parquet5,31932pgs_id, ppm_id, pss_id, pgp_id, trait_reported, covariates, pmid, doi, ... (32 total)
metadata/publications.parquet7908pgp_id, first_author, pmid, doi, title, authors, journal, date_publication
metadata/pgs_quality_scores.parquet60446pgs_id, trait, trait_grouped, trait_efo_id, synthetic_score, tier, auroc, cindex, ... (46 total)

Scoring files (data/scores/)

Individual scoring weight files, one parquet per PGS ID. Each file contains variant-level effect weights with harmonised positions.

InfoValue
Total files5,337
Naming pattern{PGS_ID}_hmPOS_GRCh38.parquet
Compressionzstd (level 9)
ExamplesPGS000001_hmPOS_GRCh38.parquet, PGS000002_hmPOS_GRCh38.parquet, PGS000003_hmPOS_GRCh38.parquet, ... (5,337 files total)

Metadata Cleanup Pipeline

The raw PGS Catalog CSVs undergo several transformations:

  1. 1.Column renaming — verbose PGS column names (e.g. Polygenic Score (PGS) ID) are mapped to short snake_case equivalents (pgs_id).
  2. 2.Genome build normalization — 9 raw build variants (hg19, hg37, hg38, NCBI36, hg18, NCBI35, GRCh37, GRCh38, NR) are mapped to canonical values: GRCh37, GRCh38, GRCh36, or NR.
  3. 3.Metric string parsing — performance metrics stored as strings like "1.55 [1.52,1.58]" or "-0.7 (0.15)" are parsed into structured numeric columns (*_estimate, *_ci_lower, *_ci_upper, *_se) for OR, HR, Beta, AUROC, and C-index.
  4. 4.Performance flattening — performance metrics are joined with evaluation sample sets to include sample size (n_individuals) and ancestry (ancestry_broad).
  5. 5.Best performance selection — best_performance.parquet contains one row per PGS ID, selecting the evaluation with the largest sample size (with a preference for European-ancestry cohorts).

Scoring File Schema

Each scoring parquet contains variant-level data from the PGS Catalog harmonised files:

  • —rsID — dbSNP rsID
  • —chr_name / chr_position — chromosome and position (original)
  • —effect_allele / other_allele — alleles
  • —effect_weight — variant weight (beta, log-OR, etc.)
  • —hm_chr / hm_pos — harmonised chromosome and position (GRCh38)
  • —hm_inferOtherAllele — inferred other allele during harmonisation
  • —PGS header metadata is embedded as file-level Parquet metadata under the key pgs_catalog_header.
  • —In data/metadata/scores.parquet, ftp_link points to the parquet mirror in this dataset (data/scores/...parquet) for fast parquet-first loading; original EBI links are preserved in ftp_link_ebi.

Usage

Load metadata with polars

python
import polars as pl
from huggingface_hub import hf_hub_download

path = hf_hub_download(
    repo_id="just-dna-seq/pgs-catalog",
    filename="data/metadata/scores.parquet",
    repo_type="dataset",
)
scores = pl.read_parquet(path)
print(scores.head())

Load a scoring file

python
path = hf_hub_download(
    repo_id="just-dna-seq/pgs-catalog",
    filename="data/scores/PGS000001_hmPOS_GRCh38.parquet",
    repo_type="dataset",
)
weights = pl.read_parquet(path)
print(weights.head())

Source & License

  • —Source: PGS Catalog (EBI / NHGRI)
  • —License: CC BY 4.0
  • —Citation: Lambert, S.A. et al. The Polygenic Score Catalog as an open database for reproducibility and systematic evaluation. Nature Genetics 53, 420-425 (2021). doi:10.1038/s41588-021-00783-5