kozo2/metabolomics-edge-node2vec
Node2Vec embeddings for cross-study metabolomics co-response graphs
128-dimensional Node2Vec embeddings for two undirected graphs of metabolite co-response edges that pass a chi-square validity filter (every expected frequency ≥ 5). Held-out link prediction reaches AUC 0.932 and 0.928, against degree-only baselines of 0.826 and 0.888.
The graphs, node properties and full pipeline are in the companion dataset repository: kozo2/metabolomics-edges-expected-ge5.
Files
Each file records the exact hyperparameters that produced it under args, and node_id[i] is the original study/assay/feature ID of embedding row i.
Using it
import torch
ck = torch.load("edge_MLvsMW_filtered_expected_ge5_n2v.pt", weights_only=False)
z = torch.nn.functional.normalize(ck["embedding"], dim=1) # cosine space
idx = {node: i for i, node in enumerate(ck["node_id"])}
q = z[idx["MTBLS311_0001_00001537"]]
top = (z @ q).topk(11).indices[1:] # drop self
print([ck["node_id"][i] for i in top])Scores are cosine similarities — that is the metric the embeddings were trained under random walks and evaluated with, so use normalized vectors rather than raw dot products.
Training
Trained on one NVIDIA H100 NVL. sparse=True makes the embedding table the only parameter tensor, which is why the optimizer is SparseAdam.
Evaluation
heldout_check.py deduplicates undirected edges, holds out a random 10%, retrains from scratch on the remaining 90%, then scores held-out edges against an equal number of sampled non-edges by cosine similarity. The control is the preferential-attachment baseline d_u · d_v on training degrees, which answers whether the embedding learned anything beyond "popular nodes connect".
Species neighbour purity — of each node's 10 nearest embeddings by cosine, the fraction sharing its species label — is 48.8% overall for edge_ML and 59.2% for edge_MLvsMW, rising to 75–79% for the dominant host species. For the bipartite edge_MLvsMW graph, the same measure on source database gives 67.5% overall (75.4% ST, 47.7% MTBLS).
Caveats
- Training budget matters more than it looks. At 20 epochs,
edge_MLscored a held-out AUC of 0.791 — below its own 0.826 degree baseline. The embeddings only beat preferential attachment once training converges.--epochsdefaults to 200 for this reason; do not shorten it without re-runningheldout_check.py. - On
edge_MLvsMWthe margin over the degree baseline is thin (+0.040). That graph is bipartite with a few very high-degree studies, so much of its link structure is explained by degree alone. edge_MLvsMWloss is flat from about epoch 19, so 200 epochs is more budget than that graph needs.- Embeddings are transductive: there is no way to embed a node that was not in the training graph without retraining.
