gitmodelmujtaba/sapbert-snomed-loinc-rxnorm
Clinical SapBERT Tri-Linker: Multi-Ontology Linking & Graph-RAG Engine
     
A specialized biomedical representation model, supervised Cross-Encoder Reranker, and Contrastive Metric Projection Adapter engineered for multi-ontology clinical entity resolution across:
- SNOMED CT: Clinical findings, disorders, procedures, and body structures (638,238 concepts).
- RxNorm: Active ingredients, branded formulations, and dosages (316,330 concepts).
- LOINC: Laboratory tests, observations, and diagnostic measurements (287,811 concepts).
Total Unified Vocabulary: Over 1,242,379 standardized clinical concepts.
๐ Key Updates: Active Learning & 2.24M-Edge Graph-RAG
- Hard-Negative Contrastive Adapter: Trained on 31,505 mined hard-negative triplets from EHR gold annotations, driving triplet margin loss from
0.1352down to `0.1109` to aggressively separate confusing semantic neighbors. - 1,843 Curated Concept Overrides: Embedded dictionary mapping high-acuity medical shorthand, abbreviations, and misspellings directly to standard SNOMED CT and RxNorm identifiers.
- 2.24M Clinical Relation Graph & Graph-RAG Engine: Integrated with an indexed clinical knowledge graph containing 2,248,500 directed relations (
treats,caused_by,indicated_for,evaluates,anatomical_site), enabling sub-millisecond 1-hop lookups and multi-hop pathway discovery.
๐ Quickstart: Embedding & Similarity
import torch
import torch.nn.functional as F
from transformers import AutoTokenizer, AutoModel
# 1. Load model and tokenizer
repo_id = "gitmodelmujtaba/sapbert-snomed-loinc-rxnorm"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModel.from_pretrained(repo_id)
model.eval()
# 2. Clinical mentions vs Standard Ontology Concepts
mentions = [
"heart attack",
"lap chole",
"sugar in urine",
]
concepts = [
"Acute myocardial infarction (disorder)",
"Laparoscopic cholecystectomy (procedure)",
"Glycosuria (finding)",
]
# 3. Generate 768-dim CLS embeddings
def get_embeddings(texts):
inputs = tokenizer(texts, padding=True, truncation=True, return_tensors="pt")
with torch.no_grad():
outputs = model(**inputs)
return F.normalize(outputs.last_hidden_state[:, 0, :], p=2, dim=1)
m_emb = get_embeddings(mentions)
c_emb = get_embeddings(concepts)
# 4. Cosine similarity matrix
similarity = torch.mm(m_emb, c_emb.T)
for i, mention in enumerate(mentions):
best_idx = similarity[i].argmax().item()
print(f"'{mention}' --> '{concepts[best_idx]}' (similarity: {similarity[i][best_idx]:.4f})")๐ Graph-RAG Multi-Hop Querying
The model's concepts interface directly with the 2.24M clinical relation graph for multi-hop clinical pathway discovery:
[1-Hop] laparoscopic cholecystectomy --[treats]--> gallstone pancreatitis (score: 0.88)
[2-Hop] laparoscopic cholecystectomy -[treats]-> abdominal pain -[caused_by]-> acute pancreatitis
[3-Hop] laparoscopic cholecystectomy -[treats]-> surgical -[caused_by]-> nausea -[caused_by]-> acute pancreatitis๐ License & Attribution
- License: Apache 2.0
- Author: Mujtaba Hussain (`gitmodelmujtaba`)
- NER Partner Model: GLiNER-BioMed
- Live Space: Clinical GLiNER & SapBERT RelEC + Graph-RAG
