Team Ai
Modelpublic

kozo2/metabolomics-edge-node2vec

sourceHugging Facecc-by-4.0updated 14d agoView on Hugging Face
0likes
Model Card

Node2Vec embeddings for cross-study metabolomics co-response graphs

128-dimensional Node2Vec embeddings for two undirected graphs of metabolite co-response edges that pass a chi-square validity filter (every expected frequency ≥ 5). Held-out link prediction reaches AUC 0.932 and 0.928, against degree-only baselines of 0.826 and 0.888.

The graphs, node properties and full pipeline are in the companion dataset repository: kozo2/metabolomics-edges-expected-ge5.

Files

FileContentsSize
edge_ML_filtered_expected_ge5_n2v.pt{embedding: [848, 128] float32, node_id: [848], args: {...}}453 KB
edge_MLvsMW_filtered_expected_ge5_n2v.pt{embedding: [15758, 128] float32, node_id: [15758], args: {...}}8.6 MB
node2vec_model.py, config.pyModel definition, training loop, embedding export—
heldout_check.pyThe held-out link-prediction evaluation reported below—

Each file records the exact hyperparameters that produced it under args, and node_id[i] is the original study/assay/feature ID of embedding row i.

Using it

python
import torch

ck = torch.load("edge_MLvsMW_filtered_expected_ge5_n2v.pt", weights_only=False)
z = torch.nn.functional.normalize(ck["embedding"], dim=1)   # cosine space
idx = {node: i for i, node in enumerate(ck["node_id"])}

q = z[idx["MTBLS311_0001_00001537"]]
top = (z @ q).topk(11).indices[1:]                          # drop self
print([ck["node_id"][i] for i in top])

Scores are cosine similarities — that is the metric the embeddings were trained under random walks and evaluated with, so use normalized vectors rather than raw dot products.

Training

`edge_ML``edge_MLvsMW`
graph848 nodes, 3,697 edges15,758 nodes, 89,277 edges
parameters108,5442,017,024
embedding_dim128128
walk_length / context_size20 / 1020 / 10
walks_per_node / negatives10 / 110 / 1
p / q1.0 / 1.0 (unbiased)1.0 / 1.0 (unbiased)
batch_size / lr / optimizer128 / 0.01 / SparseAdam128 / 0.01 / SparseAdam
epochs / wall time200 / 31 s200 / 3 m 05 s
loss, first → last9.02 → 0.8156.73 → 0.876

Trained on one NVIDIA H100 NVL. sparse=True makes the embedding table the only parameter tensor, which is why the optimizer is SparseAdam.

Evaluation

heldout_check.py deduplicates undirected edges, holds out a random 10%, retrains from scratch on the remaining 90%, then scores held-out edges against an equal number of sampled non-edges by cosine similarity. The control is the preferential-attachment baseline d_u · d_v on training degrees, which answers whether the embedding learned anything beyond "popular nodes connect".

`edge_ML``edge_MLvsMW`
train / held-out edges3,328 / 36980,350 / 8,927
nodes isolated by the split42413
held-out AUC (never seen)0.9320.928
train AUC (in-sample)0.9850.985
degree baseline (held-out)0.8260.888
margin over baseline+0.106+0.040

Species neighbour purity — of each node's 10 nearest embeddings by cosine, the fraction sharing its species label — is 48.8% overall for edge_ML and 59.2% for edge_MLvsMW, rising to 75–79% for the dominant host species. For the bipartite edge_MLvsMW graph, the same measure on source database gives 67.5% overall (75.4% ST, 47.7% MTBLS).

Caveats

  • —Training budget matters more than it looks. At 20 epochs, edge_ML scored a held-out AUC of 0.791 — below its own 0.826 degree baseline. The embeddings only beat preferential attachment once training converges. --epochs defaults to 200 for this reason; do not shorten it without re-running heldout_check.py.
  • —On edge_MLvsMW the margin over the degree baseline is thin (+0.040). That graph is bipartite with a few very high-degree studies, so much of its link structure is explained by degree alone.
  • —edge_MLvsMW loss is flat from about epoch 19, so 200 epochs is more budget than that graph needs.
  • —Embeddings are transductive: there is no way to embed a node that was not in the training graph without retraining.