Team Ai
Datasetpublic

longevity-genie/atomica_longevity_proteins

Longevity Protein Structures Dataset Quick Start # Load the dataset index from datasets import load_dataset ds = load_dataset("longevity-genie/atomica_longevity_proteins") # Or use Polars directly import polars as pl df = pl.read_parquet("hf://datasets/longevity-genie/atomica_longevity_proteins/atomica_index.parquet") # Find KEAP1 structures keap1 = df.filter(pl.col("gene_symbols").list.contains("KEAP1")) print(f"Found {len(keap1)} KEAP1 structures")… See the full description on the dataset page: https://huggingface.co/datasets/longevity-genie/atomica_longevity_proteins.

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes150downloads
Dataset Card

Longevity Protein Structures Dataset

![License: MIT](https://opensource.org/licenses/MIT) ![Dataset Size](https://huggingface.co/datasets/longevity-genie/atomicalongevityproteins) ![ATOMICA](https://github.com/longevity-genie/ATOMICA)

Quick Start

python
# Load the dataset index
from datasets import load_dataset
ds = load_dataset("longevity-genie/atomica_longevity_proteins")

# Or use Polars directly
import polars as pl
df = pl.read_parquet("hf://datasets/longevity-genie/atomica_longevity_proteins/atomica_index.parquet")

# Find KEAP1 structures
keap1 = df.filter(pl.col("gene_symbols").list.contains("KEAP1"))
print(f"Found {len(keap1)} KEAP1 structures")

Overview

This dataset contains comprehensive structural analysis of key longevity-related proteins using the ATOMICA deep learning model. The dataset includes 94 protein structures spanning five major protein families involved in oxidative stress response, pluripotency, and lipid metabolism pathways associated with aging and longevity.

Dataset Creation

Analysis Tool: ATOMICA - A pretrained deep learning model for protein structure interaction scoring Model Architecture:

  • —Hidden size: 32
  • —Edge size: 32
  • —K-neighbors: 8
  • —Layers: 4
  • —Global message passing enabled
  • —Fragmentation method: PS_300

Processing Environment:

  • —Device: CUDA GPU
  • —Average GPU Memory per structure: ~2.0 GB
  • —Average processing time: ~44 seconds per structure

Protein Families Included

1. NRF2-KEAP1 System (Oxidative Stress Response)

  • —NRF2 (NFE2L2): 19 structures
  • —KEAP1: 47 structures
  • —Focus: KEAP1 mutations in Neoaves, SKN-1 lifespan effects in C. elegans

2. SOX2 (Pluripotency Factor)

  • —SOX2: 8 structures
  • —Focus: SuperSOX modifications for enhanced reprogramming

3. APOE (Lipid Metabolism & Alzheimer's Risk)

  • —APOE variants (E2/E3/E4): 9 structures
  • —Focus: Longevity associations and Alzheimer's disease risk factors

4. OCT4 (Reprogramming Factor)

  • —OCT4 (POU5F1): 4 structures
  • —Focus: OCT6 conversion into reprogramming factor, Yamanaka factors

Total Structures: 94 high-resolution protein structures

File Structure

Dataset Index (NEW!)

The dataset now includes `atomica_index.parquet` - a comprehensive queryable index of all structures:

python
import polars as pl

# Load the index
df = pl.read_parquet("atomica_index.parquet")

# Query structures by gene
keap1_structures = df.filter(pl.col("gene_symbols").list.contains("KEAP1"))

# Find structures with high resolution
high_res = df.filter(pl.col("critical_residues_count") > 100)

# Get all human structures
human = df.filter(pl.col("organisms").list.contains("Homo sapiens"))

Index Schema:

  • —pdb_id: PDB identifier (uppercase)
  • —cif_path, metadata_path, summary_path, critical_residues_path, interact_scores_path, pymol_path: Relative file paths
  • —critical_residues_count: Number of critical residues identified
  • —total_time_seconds, gpu_memory_mb_max: Processing statistics
  • —metadata_found: Whether PDB metadata was successfully resolved
  • —title: Structure title from PDB
  • —uniprot_ids: List of UniProt identifiers
  • —organisms: List of organism names
  • —taxonomy_ids: List of NCBI taxonomy IDs
  • —gene_symbols: List of gene symbols
  • —ensembl_ids: List of Ensembl gene IDs
  • —structures_json: Detailed structure information (JSON string)

The index file is automatically detected by Hugging Face and can be explored in the Dataset Viewer tab!

Per-Structure Files

For each PDB structure (e.g., 6ht5), the dataset contains:

dataset/
├── atomica_index.parquet                 # QUERYABLE INDEX (NEW!)
└── {pdb_id}/                              # Per-structure directory
    ├── {pdb_id}.cif                      # Original structure file (mmCIF format)
    ├── {pdb_id}_metadata.json            # PDB metadata (if available)
    ├── {pdb_id}_interact_scores.json     # ATOMICA interaction scores
    ├── {pdb_id}_summary.json             # Processing summary statistics
    ├── {pdb_id}_critical_residues.tsv    # Ranked critical residues
    └── {pdb_id}_pymol_commands.pml       # PyMOL visualization script

File Descriptions

1. {pdb_id}.cif

  • —Format: mmCIF (Macromolecular Crystallographic Information File)
  • —Content: 3D atomic coordinates, experimental data, and metadata
  • —Source: RCSB Protein Data Bank
  • —Size: ~100-500 KB per structure

2. {pdb_id}_metadata.json

  • —Format: JSON
  • —Content: PDB metadata including:
  • —Structure resolution
  • —Experimental method (X-ray, NMR, Cryo-EM)
  • —Authors and publication info
  • —Organism source
  • —Note: May contain error if SSL certificate verification fails

3. {pdb_id}_interact_scores.json

  • —Format: JSON
  • —Content: ATOMICA deep learning predictions
  • —id: Structure identifier
  • —cos_distances: Array of cosine distance scores (one per residue block)
  • —block_idx: Sequential indices for residue blocks (1-N)
  • —time_seconds: Processing time
  • —peak_memory_mb: GPU memory usage

Score Interpretation:

  • —Scores range from 0 to 1 (typically 0.999+)
  • —Higher scores = more structurally critical residues
  • —Lower scores = potential mutation hotspots or flexibility regions

4. {pdb_id}_summary.json

  • —Format: JSON
  • —Content: Processing statistics
  • —Total processing time
  • —GPU memory usage (mean, max)
  • —Time per structure statistics
  • —Input/output file paths
  • —Device information

5. {pdb_id}_critical_residues.tsv

  • —Format: Tab-separated values (TSV)
  • —Content: Ranked list of structurally critical residues
  • —Rank (1-N)
  • —Residue name (3-letter code)
  • —Chain ID
  • —Position in sequence
  • —ATOMICA_SCORE (cosine distance)
  • —Importance Delta (deviation from mean)

Example:

Rank   Residue    Chain   Position   ATOMICA_SCORE   Importance Delta
----------------------------------------------------------------------
1      GLY171     E       171        0.999856        0.0144%
2      SER216     E       216        0.999867        0.0133%
3      ILE215     E       215        0.999875        0.0125%

6. {pdb_id}_pymol_commands.pml

  • —Format: PyMOL script
  • —Content: Commands for visualizing critical residues
  • —Usage:
bash
  pymol dataset/{pdb_id}.cif
  @dataset/{pdb_id}_pymol_commands.pml
  • —Features: Color-coded residues by criticality score

ATOMICA Scores Explained

Cosine Distance Scoring

ATOMICA uses a graph neural network to predict residue-level interaction importance:

  1. 1.Input: 3D protein structure (atomic coordinates)
  2. 2.Processing: Graph representation with residues as nodes
  3. 3.Output: Cosine distance scores per residue block

Score Ranges

  • —0.9999+ (Critical): Essential structural residues
  • —Active site residues
  • —Binding interface residues
  • —Structural core residues
  • —0.9998-0.9999 (Important): Functionally relevant residues
  • —Secondary binding sites
  • —Conformational switches
  • —<0.9998 (Variable): Potential mutation sites
  • —Surface loops
  • —Flexible regions
  • —Potential engineering targets

Importance Delta

  • —Calculated as: (max_score - residue_score) / max_score × 100%
  • —Lower delta = more critical residue
  • —Top 10 residues typically have delta <0.02%

Use Cases

0. Query the Index (NEW!)

The atomica_index.parquet enables powerful filtering and analysis:

python
import polars as pl

# Load index
df = pl.read_parquet("atomica_index.parquet")

# Find all NRF2 structures
nrf2 = df.filter(pl.col("gene_symbols").list.contains("NFE2L2"))

# Get KEAP1-NRF2 complexes
complexes = df.filter(
    pl.col("gene_symbols").list.contains("KEAP1") & 
    pl.col("gene_symbols").list.contains("NFE2L2")
)

# Find structures by UniProt ID
my_protein = df.filter(pl.col("uniprot_ids").list.contains("Q14145"))

# Get statistics
print(f"Total structures: {len(df)}")
print(f"Unique genes: {len(set(gene for genes in df['gene_symbols'] for gene in genes))}")
print(f"Organisms: {set(org for orgs in df['organisms'] for org in orgs)}")

1. Mutation Impact Prediction

Identify residues where mutations would likely destabilize the protein:

python
# Load critical residues
import pandas as pd
critical = pd.read_csv('dataset/6ht5_critical_residues.tsv', sep='\t')
# Top 10 most critical = avoid mutations
avoid_mutations = critical.head(10)

2. Protein Engineering

Target residues with lower scores for:

  • —Enhanced stability variants
  • —Altered binding specificity
  • —Improved expression

3. Drug Design

Identify binding pockets and interaction hotspots:

  • —Critical residues in binding interfaces
  • —Allosteric regulation sites
  • —Druggable pockets

4. Comparative Analysis

Compare residue criticality across:

  • —APOE variants (E2 vs E3 vs E4)
  • —NRF2-KEAP1 mutations in different species
  • —OCT4 vs OCT6 structural differences

5. MCP Server Integration (NEW!)

Use the atomica-mcp server for programmatic access:

bash
# Install the MCP server
pip install atomica-mcp

# Use in Python
from atomica_mcp import AtomicaMCP

# Query structures by gene
results = mcp.atomica_search_by_gene("KEAP1", species="Homo sapiens")

# Get structure files for a PDB ID
files = mcp.atomica_get_structure_files("6ht5")

# Search by UniProt ID
structures = mcp.atomica_search_by_uniprot("Q14145")

The MCP server provides:

  • —Fast local index queries (instant)
  • —Automatic dataset download and management
  • —Integration with Claude Desktop and other AI tools
  • —External PDB/UniProt API fallback for comprehensive searches

Dataset Statistics

Protein FamilyStructuresAvg ResolutionMethods
NRF2191.5-2.8 ÅX-ray, NMR
KEAP1471.35-3.5 ÅX-ray, Cryo-EM
SOX282.3-5.1 ÅX-ray, NMR, Cryo-EM
APOE92.0-2.5 Å, NMRX-ray, NMR
OCT44Cryo-EMX-ray, Cryo-EM

Total Dataset Size: ~40-50 MB (compressed) Total Residues Analyzed: ~15,000-20,000

Data Quality

Structure Quality Metrics

  • —X-ray structures: 1.35-3.5 Å resolution
  • —NMR structures: Ensemble of 10-20 models
  • —Cryo-EM structures: 3.0-5.1 Å resolution

ATOMICA Processing Quality

  • —GPU Memory: Stable ~2 GB per structure
  • —Processing Time: Consistent ~44s per structure
  • —Score Convergence: All scores >0.998 (high confidence)

Citation

If you use this dataset, please cite:

  1. 1.ATOMICA Model:
   ATOMICA: A deep learning model for protein structure analysis
   GitHub: https://github.com/longevity-genie/ATOMICA
  1. 1.PDB Structures:
   Berman, H.M. et al. (2000) The Protein Data Bank. 
   Nucleic Acids Research, 28: 235-242.
  1. 1.Specific protein references: See individual PDB entries for citations

Technical Notes

Requirements for Reproducing Analysis

bash
# Install ATOMICA
git clone https://github.com/longevity-genie/ATOMICA
cd ATOMICA

# Download PDB structures
uv run pdb download <pdb_id>

# Run analysis
uv run interact-score --input downloads/pdbs/<pdb_id>.cif

Known Limitations

  1. 1.Metadata Retrieval: Some structures may have SSL certificate errors
  2. 2.Resolution Dependent: Lower resolution structures (<3 Å) have more uncertainty
  3. 3.NMR Structures: Multiple models may give variable scores
  4. 4.Fragment Structures: Incomplete proteins analyzed as-is

Future Extensions

Potential additions to this dataset:

  • —SKN-1 structures from C. elegans (when available)
  • —OCT6 structures for comparative analysis
  • —Additional APOE variant structures
  • —Molecular dynamics trajectories
  • —Ligand-bound complexes

PDB Structures for Longevity-Related Proteins

NRF2 (NFE2L2) - Nuclear Factor Erythroid 2-Related Factor 2

PDB IDSpeciesLigand/ComplexResolutionComments
2FLUHumanKEAP1 complex1.50 ÅNRF2 peptide (69-84) with KEAP1, X-ray
4IFLHumanKEAP1 complex1.80 ÅNRF2 peptide with KEAP1 Kelch domain
5WFVHumanKEAP1 complex1.91 ÅNRF2 peptide (76-84) with KEAP1
6T7VHumanKEAP1 complex2.60 ÅNRF2 peptide with KEAP1, X-ray
7K28HumanKEAP1 complex2.15 ÅNRF2 peptide (77-84)
7K29HumanKEAP1 complex2.20 ÅNRF2 peptide (76-84)
7K2AHumanKEAP1 complex1.90 ÅNRF2 peptide (76-83)
7K2BHumanKEAP1 complex2.31 ÅNRF2 peptide (77-83)
7K2CHumanKEAP1 complex2.11 ÅNRF2 peptide (77-82)
7K2DHumanKEAP1 complex2.21 ÅNRF2 peptide (77-82)
7K2EHumanKEAP1 complex2.03 ÅNRF2 peptide (77-82)
7K2KHumanKEAP1 complex1.98 ÅNRF2 peptide (77-82)
7O7BHumanApoNMRNeh1 domain (445-523), solution structure
7X5EHumanMAFG complex2.30 ÅbZIP domain (452-560) with MAFG
7X5FHumanMAFG complex2.60 ÅbZIP domain (452-560) with MAFG
7X5GHumanMAFG complex2.30 ÅbZIP domain (452-560) with MAFG
3ZGCHumanKEAP1 complex2.20 ÅWith KEAP1 Kelch domain
8EJRHumanKEAP1 complex2.08 ÅNRF2 peptide with KEAP1
8EJSHumanKEAP1 complex2.82 ÅNRF2 peptide with KEAP1

KEAP1 - Kelch-like ECH-Associated Protein 1

PDB IDSpeciesLigand/ComplexResolutionComments
6LRZHumanApo1.54 ÅVery high resolution, residues 311-616
1ZGKHumanApo1.35 ÅKelch domain (321-609), highest resolution
6HWSHumanApo1.75 ÅKelch domain (321-609)
1U6DHumanApo1.85 ÅKelch domain (321-609)
2FLUHumanNRF2 peptide1.50 ÅComplex with NRF2 (69-84)
4IFLHumanNRF2 peptide1.80 ÅKelch domain with NRF2
5WFVHumanNRF2 peptide1.91 ÅKelch domain (320-612)
7K28-7K2KHumanNRF2 peptides1.98-2.31 ÅSeries of NRF2 binding studies
4CXIHumanCDDO ligand2.35 ÅBTB domain (48-180) with inhibitor
4CXJHumanCDDO ligand2.80 ÅBTB domain (48-180)
4CXTHumanLigand2.66 ÅBTB domain (48-180)
5DADHumanApo2.61 ÅBTB domain (49-182)
5DAFHumanLigand2.37 ÅBTB domain (49-182)
5GITHumanLigand2.19 ÅBTB domain (48-180)
5NLBHumanCUL3 complex3.45 ÅBACK domain (51-204) with CUL3
5F72HumanPeptide1.85 ÅFull Kelch domain (321-611)
3VNGHumanPeptide2.10 ÅKelch domain (321-609)
3VNHHumanPeptide2.10 ÅKelch domain (321-609)
3ZGCHumanNRF2 peptide2.20 ÅKelch domain complex
3ZGDHumanPeptide1.98 ÅKelch domain (321-609)
4IFJHumanPeptide1.80 ÅKelch domain (321-609)
4IFNHumanPeptide2.40 ÅKelch domain (321-609)
4IQKHumanPeptide1.97 ÅKelch domain (321-609)
4IN4HumanLigand2.59 ÅKelch domain (321-609)
4L7BHumanLigand2.41 ÅKelch domain (321-609)
4L7CHumanLigand2.40 ÅKelch domain (321-609)
4L7DHumanLigand2.25 ÅKelch domain (321-609)
4N1BHumanLigand2.55 ÅKelch domain (321-609)
4XMBHumanPeptide2.43 ÅKelch domain (321-609)
5WFLHumanPeptide1.93 ÅKelch domain (312-624)
5WG1HumanPeptide2.02 ÅKelch domain (320-612)
5WHLHumanPeptide2.50 ÅKelch domain (312-624)
5WHOHumanPeptide2.23 ÅKelch domain (312-624)
5WIYHumanPeptide2.23 ÅKelch domain (312-624)
5X54HumanPeptide2.30 ÅKelch domain (321-609)
6FFMHumanLigand2.20 ÅBTB domain (48-180)
6FMPHumanPeptide2.92 ÅKelch domain (321-609)
6FMQHumanPeptide2.10 ÅKelch domain (321-609)
6ROGHumanPeptide2.16 ÅKelch domain (321-609)
6SP1HumanPeptide2.57 ÅKelch domain (321-609)
6SP4HumanPeptide2.59 ÅKelch domain (321-609)
6T7ZHumanPeptide2.00 ÅKelch domain (321-609)
6TG8HumanPeptide2.75 ÅKelch domain (322-609)
7EXIHumanApoMultipleRecent full-length BTB-BACK-Kelch
7X4WHumanApoMultipleBTB domain structure
7X4XHumanApoMultipleBTB domain structure

SOX2 - Sex-Determining Region Y-Box 2

PDB IDSpeciesLigand/ComplexResolutionComments
1O4XHumanDNANMRHMG domain (39-121), solution structure
2LE4HumanApoNMRHMG domain (39-118), solution structure
6WX8HumanDNA complex2.30 ÅHMG domain (39-127), best resolution X-ray
6WX7HumanDNA complex2.70 ÅHMG domain (39-127) with DNA
6WX9HumanDNA complex2.80 ÅHMG domain (39-127) with DNA
6T90HumanComplex3.05 ÅCryo-EM structure (37-118)
6YOVHumanComplex3.42 ÅCryo-EM structure (37-118)
6T7BHumanComplex5.10 ÅCryo-EM structure (36-121)

APOE - Apolipoprotein E (variants E2, E3, E4)

APOE2 Variant

PDB IDSpeciesLigand/ComplexResolutionComments
1LE2HumanApoX-rayAPOE2 N-terminal domain, reduced receptor binding

APOE3 Variant (Wild-type, neutral)

PDB IDSpeciesLigand/ComplexResolutionComments
1LPEHumanApo2.50 ÅLDL receptor-binding domain, four-helix bundle
1NFNHumanApoX-rayAPOE3 N-terminal domain structure
2L7BHumanApoNMRFull-length APOE3, solution structure

APOE4 Variant (Alzheimer's risk factor)

PDB IDSpeciesLigand/ComplexResolutionComments
1B68HumanHeparin octasaccharideX-rayAPOE4 22K fragment (1-191), disease-relevant
1LE4HumanApo2.50 ÅAPOE4 N-terminal, E112R mutation, altered function
8AX8HumanApoX-rayAPOE4 N-terminal, aggregation-prone conformation

Other APOE Structures

PDB IDSpeciesLigand/ComplexResolutionComments
1OEFHumanSDS micelleNMRC-terminal peptide (263-286), lipid binding
1YA9MouseApo2.09 ÅMouse APOE N-terminal (wild-type)

OCT4 (POU5F1) - Octamer-Binding Transcription Factor 4

PDB IDSpeciesLigand/ComplexResolutionComments
3L1PMouseDNAX-rayOct4 POU domain (131-282), reprogramming factor
8G86HumanNucleosome/nMatn1 DNACryo-EMOCT4 bound to chromatin, pioneer factor activity
8G87HumanNucleosome/nMatn1 DNACryo-EMFocused refinement of OCT4-nucleosome
6HT5Human/MouseSOX2-UTF1-DNAX-rayOct4-Sox2 complex on DNA (referenced)

Notes and Research Context

NRF2-KEAP1 System

  • —Oxidative Stress Response: NRF2 is sequestered by KEAP1 under normal conditions
  • —KEAP1 Mutations in Neoaves: Structures show binding interface mutations
  • —Drug Development: Multiple inhibitor-bound KEAP1 structures available
  • —Key Interface: ETGE and DLG motifs of NRF2 bind to Kelch domain of KEAP1

SKN-1 in C. elegans

  • —No direct human homolog structures, but functionally similar to NRF2
  • —Studies show lifespan extension effects through oxidative stress pathways

SOX2 Modifications

  • —SuperSOX variants: Modifications that enhance reprogramming efficiency
  • —Structures show HMG box DNA-binding mechanism
  • —Works synergistically with OCT4 in pluripotency

APOE Variants and Longevity

  • —APOE2: Protective, associated with longevity (Cys112, Cys158)
  • —APOE3: Neutral variant (Cys112, Arg158)
  • —APOE4: Risk factor for Alzheimer's (Arg112, Arg158)
  • —Structural differences affect lipid binding and receptor interactions

OCT4-OCT6 Conversion

  • —OCT4 (POU5F1): Pluripotency factor, one of Yamanaka factors
  • —OCT6 (POU3F1): Neuronal POU factor
  • —Converting OCT6 to reprogramming function requires understanding POU domain specificity
  • —Limited OCT6 structures available; comparative modeling with OCT4 recommended

Resolution Quality Guide

  • —< 2.0 Å: Atomic detail, excellent for drug design
  • —2.0-2.5 Å: High quality, suitable for most analyses
  • —2.5-3.5 Å: Good quality, suitable for overall structure
  • —> 3.5 Å or NMR: Domain organization, dynamics studies
  • —Cryo-EM: Large complexes, native-like states

License

  • —PDB Structures: Public domain (RCSB PDB)
  • —ATOMICA Predictions: Check ATOMICA repository for license
  • —Dataset Compilation: [Your license here]

Contact

For questions about:


Last Updated: October 2025 Dataset Version: 1.0 Total Structures: 94 Analyzed Using: ATOMICA v1.0 Data Sources: RCSB Protein Data Bank, UniProt, Literature searches