Indhu27/MachineLearning_Algorithms
0
1import streamlit as st2 3st.set_page_config(page_title="svm", page_icon="🤖", layout="wide")4st.markdown("""5 <style>6 .stApp {7 background-color: #4A90E2;8 }9 h1, h2, h3 {10 color: #003366; /* Adjust this if needed to match your background */11 }12 .custom-font, p {13 font-family: 'Arial', sans-serif;14 font-size: 18px;15 color: white; /* Making all inside text white */16 line-height: 1.6;17 }18 </style>19 """, unsafe_allow_html=True)20 21 22# Title of the Streamlit Application23st.markdown("<h1 style='color: #003366;'>Support Vector Machines (SVM) in Machine Learning</h1>",unsafe_allow_html=True)24 25 26def main():27 28 st.write("""29 Support Vector Machines (SVM) is a **supervised learning algorithm** used for both **classification** and **regression** problems. 30 However, it is mostly used for classification tasks in real-world applications.31 32 SVM is a **parametric model** and is also known as a **linear model** in its basic form. It works by finding the optimal decision 33 boundary that maximizes the margin between different classes.34 """)35 st.image("svm.png",width=700)36 37 st.subheader("Types of SVM")38 st.write("""39 1. **Support Vector Classifier (SVC)** - Used for classification problems.40 2. **Support Vector Regression (SVR)** - Used for regression problems.41 """)42 43 st.markdown("</h2 style='color:'#003366;'>SVC: Support Vector Classifier</h2>",unsafe_allow_html=True)44 st.subheader("Working of SVC")45 st.write("""46 1. Randomly initialize weights and draw a line (hyperplane) to separate classes.47 2. Draw parallel lines (support vectors) at equal distances until one of the lines touches a data point.48 3. The **margin** is calculated as the distance between these support vectors.49 4. The objective is to **maximize this margin** while minimizing misclassification.50 """)51 52 st.markdown("</h2 style='color:'#003366;'>Hard Margin vs. Soft Margin SVC</h2>",unsafe_allow_html=True)53 st.write("""54 - **Hard Margin SVC**: Assumes data is **perfectly linearly separable** and does not allow misclassification.55 - **Soft Margin SVC**: Allows **some misclassification** to improve generalization on new data.56 """)57 st.image("soft vs hard.png",width=700)58 59 st.subheader("Mathematical Formulation")60 st.write("""61 - **Hard Margin Condition**: Ensures all points are correctly classified and lie outside the margin.62 """)63 st.latex(r" y_i (w^T x_i + b) \geq 1")64 65 st.write("""66 - **Soft Margin Condition**: Introduces a slack variable to allow misclassification.67 """)68 st.latex(r"y_i (w^T x_i + b) \geq 1 - \xi_i")69 70 st.write("""71 ### Interpretation of Slack Variable \( \xi \)72 - \( \xi_i = 0 \) : Correct classification, point lies outside the margin.73 - \( 0 < \xi_i \leq 1 \) : Correct classification, but the point is inside the margin.74 - \( \xi_i > 1 \) : Misclassification occurs.75 """)76 77 78 st.subheader("Advantages & Disadvantages")79 st.write("""80 **Advantages:**81 - Effective in high-dimensional spaces.82 - Works well with both linearly and non-linearly separable data (using kernels).83 - Robust against overfitting in high-dimensional datasets.84 85 **Disadvantages:**86 - Computationally expensive for large datasets.87 - Requires careful tuning of hyperparameters (C, kernel type).88 """)89 90 st.markdown("</h2 style='color:'#003366;'>Dual form of SVM</h2>",unsafe_allow_html=True)91 st.write("""92 When data is **not linearly separable**, we use the **Kernel Trick** to transform the data into a higher-dimensional space where 93 it becomes separable.94 95 **Types of Kernels:**96 - **Linear Kernel**: Used when data is already linearly separable.97 - **Polynomial Kernel**: Maps data into a polynomial feature space.98 - **Radial Basis Function (RBF) Kernel**: Captures complex relationships.99 - **Sigmoid Kernel**: Mimics the behavior of a neural network.100 """)101 st.image("dualform.png",width=700)102 103 st.header("Hyperparameter Tuning in SVM")104 st.write("""105 - **C Parameter**: Controls trade-off between maximizing margin and minimizing classification error.106 - **High C**: Less misclassification, smaller margin (risk of overfitting).107 - **Low C**: More misclassification, larger margin (better generalization).108 - **Gamma (for RBF kernel)**: Determines the influence of individual training points.109 - **High Gamma**: Each point has more influence (risk of overfitting).110 - **Low Gamma**: Each point has less influence (risk of underfitting).111 """)112 113 114if __name__ == "__main__":115 main()116 