neuroknowai/binary-classification-drug-reactions
binary-classification-drug-reactions
binary-classification-drug-reactions is NeuroKnow AI's 184.4M-parameter DeBERTa-v3-based binary classifier for assessing the severity of adverse drug reactions from patient-reported narrative text. It classifies each input as Severe or Non-Severe for pharmacovigilance and drug-safety workflows.
What it does
The model classifies patient-reported drug-experience narratives into:
Non-SevereSevere
It supports adverse drug reaction triage, pharmacovigilance analysis, drug-safety workflows, biomedical text classification, and research-oriented decision-support pipelines.
Model architecture
- Base architecture:
microsoft/deberta-v3-base - Transformers architecture:
DebertaV2ForSequenceClassification - Architecture family: DeBERTa-v3
- Task: binary sequence classification
- Classes:
Non-Severe,Severe - Transformer layers: 12
- Hidden size: 768
- Attention heads: 12
- Intermediate size: 3072
- Classifier output dimension: 2
- Parameter count: 184,423,682
- Model dtype: float32
- Tokenizer: SentencePiece Unigram via
DebertaV2Tokenizer - Vocabulary size: 128,100
- Application maximum sequence length: 256 tokens
- Maximum positional embeddings: 512
Training data
The model was trained on:
- 8,153 patient-reported drug-experience narratives
- Adverse drug reaction severity labels
- A near-balanced class distribution:
- 53.4% Severe
- 46.6% Non-Severe
Additional biomedical training
The model was further adapted on OpenMed/DDI-Corpus-Processed for biomedical domain adaptation. This additional training was used to strengthen:
- Drug-interaction representations
- Pharmacological safety understanding
- Biomedical domain adaptation
- Representations relevant to adverse drug reaction severity classification
This corpus was used as additional biomedical training and is not presented as a Severe/Non-Severe adverse drug reaction dataset.
Training configuration
learning_rate: 2e-5
optimizer: adafactor
batch_size: 16
gradient_accumulation: 4
effective_batch_size: 64
epochs: 8
early_stopping_patience: 3
warmup_ratio: 0.1
lr_scheduler: cosine
weight_decay: 0.01
max_seq_length: 256
cv_folds: 5Validation metrics
Usage
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
model_id = "neuroknowai/binary-classification-drug-reactions"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
model.eval()
id2label = {
0: "Non-Severe",
1: "Severe",
}
def classify_adr(text):
inputs = tokenizer(
text,
return_tensors="pt",
truncation=True,
max_length=256,
padding=True,
)
with torch.no_grad():
logits = model(**inputs).logits
probabilities = torch.softmax(logits, dim=-1)[0]
class_id = int(probabilities.argmax())
label = id2label[class_id]
confidence = float(probabilities[class_id])
return {
"label": label,
"confidence": round(confidence, 4),
}
result = classify_adr(
"I experienced severe insomnia, heart palpitations, and extreme anxiety "
"after taking this medication for two weeks."
)
print(result)Intended use
This model is intended for research, pharmacovigilance analysis, drug-safety workflows, ADR narrative classification, and decision-support pipelines.
It should not be used as a diagnostic model or as the sole basis for clinical decisions.
Organization
Developed and maintained by NeuroKnow AI.
Hugging Face:
neuroknowai/binary-classification-drug-reactions
