morinousagi/pytorch-transformer-anomaly
Heart Disease Risk Prediction Demo Model
Disclaimer: This is a demonstration AI model for learning purposes. It does not provide medical advice.
The primary purpose of this project was to explore TabTransformer<sup>1</sup> architecture and also learn to apply TransformerEncoderLayer. This approach is typically targeted for huge complex dataset. However, small-medium sized dataset is used in this project for hands-on practice purpose only (less computationally intensive).
This project was developed through a collaborative workflow with Gemini 3 Flash. While the core architecture and debugging were AI-augmented, the final schema, transformer-based approach, logic modifications, deployment strategy and UI refinements were human-led.
Manual code review - notes:
fit_tranform() was applied on the entire dataset, prior to data split!
>>> Upon intervention, AI admitted the mistake and made correction to apply
fit_tranform() and transform() on train and test datasets respectively to prevent data leakage.About the Data
- Source: https://huggingface.co/datasets/sanadf234/Heart-Disease-Prediction-dataset
- Shape: (54897, 14)
- 1 target variable -
heart_disease - 0: Absence of heart disease
- 1: Presence of heart disease
- 12 features:
- 7 categorical:
'gender', 'chest_pain_type', 'fasting_blood_sugar', 'resting_ecg', 'exercise_angina', 'slope', 'thal' - 5 numerical:
'age', 'resting_bp', 'cholesterol', 'max_heart_rate', 'oldpeak' - Data dictionary
Data Preprocessing
- Extract & load data
- Transform data
- Encoding categorical features
- Scale numeric features
- Split data
- Save transformed data & print:
Categorical Dimensions (Embedding sizes): [2, 4, 2, 3, 2, 3, 3]
Train size: 43917 | Test size: 10980Dataset Creation
initlengetitem- Categorical features must be Long tensors for nn.Embedding
- Numerical features must be Float tensors
Model
Hybrid TabTransformer<sup>[1](https://doi.org/10.48550/arXiv.2012.06678)</sup>:
init- Embedding layer for each categorical feature - uses Encoder outputs feed into nn.Embedding
- Transformer Encoder Layer - treats categorical embeddings as a sequence of tokens
- Final MLP Head - MLP can classify non-lineraly separable classes
forwardpass- Embed each category and stack: [batch, numcats, embeddim]
- Apply Transformer Attention across the "tokens" (features)
- Flatten categorical tokens and concatenate with numerical features
Model Metrics & Evaluation
Pre-trained the model using the following hyperparameters & components:
- batch_size = 32
- epochs = 20
- lr = 0.001
- loss: BCE
- optimizer: Adam
Epoch [5/20], Loss: 0.1900
Epoch [10/20], Loss: 0.1791
Epoch [15/20], Loss: 0.1795
Epoch [20/20], Loss: 0.1769
--- Final Model Evaluation ---
AUROC Score: 0.9847
F1 Score: 0.9640
Classification Report:
precision recall f1-score support
0.0 0.92 0.80 0.85 2322
1.0 0.95 0.98 0.96 8658
accuracy 0.94 10980
macro avg 0.93 0.89 0.91 10980
weighted avg 0.94 0.94 0.94 10980
Model saved to assets/modelv1.pthEpoch [5/20], Loss: 0.1882
Epoch [10/20], Loss: 0.1816
Epoch [15/20], Loss: 0.1744
Epoch [20/20], Loss: 0.1727
--- Final Model Evaluation ---
AUROC Score: 0.9863
F1 Score: 0.9634
Classification Report:
precision recall f1-score support
0.0 0.83 0.91 0.87 2322
1.0 0.98 0.95 0.96 8658
accuracy 0.94 10980
macro avg 0.90 0.93 0.92 10980
weighted avg 0.95 0.94 0.94 10980
Model saved to assets/model.pthDeployment
- Docker & Gradio
- Demo app URL: https://huggingface.co/spaces/morinousagi/pytorch-transformer-anomaly
- Working files: https://huggingface.co/spaces/morinousagi/pytorch-transformer-anomaly/tree/main
Workflow
- Perform EDA using .ipynb
- Prepare and run ``
python src/dataset_elt.py`` --> data & asset files saved locally - Prepare ``
src/dataset.py`and`src/model.py`` - Run ``
python src/train_eval.py`` --> model.pth generated - Generate ``
Dockerfile, app.py, requirements.txt`` for Space set up - Push files to Hugging Face
Folder structure
pytorch-transformer-anomaly/
├── assets/
│ ├── model.pth # Trained weights
│ ├── model_metadata.joblib # Architecture info (cat_dims, etc.)
│ ├── num_scaler.joblib # Numerical scaler
│ └── cat_encoder.joblib # Encoded categorical features
├── data/ # Local data storage
├── src/
│ ├── dataset_prep.py # Data fetching & preprocessing
│ ├── dataset.py # PyTorch Dataset/DataLoader
│ ├── model.py # Model definition & Feature Extractor
│ └── train_eval.py # Training loop & Metrics
├── app.py # Gradio Interface
├── Dockerfile # Containerization
└── requirements.txt # Dependencies
└── README.md # Project documentationData Dictionary
Ref:
- Gemini
- [1]https://doi.org/10.48550/arXiv.2012.06678 TabTransformer: Tabular Data Modeling Using Contextual Embeddings
- https://docs.pytorch.org/docs/stable/generated/torch.nn.modules.transformer.TransformerEncoderLayer.html
