Indhu27/MachineLearning_Algorithms
0
1import streamlit as st2 3# Set page configuration4st.set_page_config(page_title="Decision Tree Theory", layout="wide")5 6# Add custom CSS for styling the headers7st.markdown("""8 <style>9 .stApp {10 background-color: #4A90E2;11 }12 h1, h2, h3 {13 color: #003366; /* Adjust this if needed to match your background */14 }15 .custom-font, p {16 font-family: 'Arial', sans-serif;17 font-size: 18px;18 color: white; /* Making all inside text white */19 line-height: 1.6;20 }21 </style>22 """, unsafe_allow_html=True)23 24# Main Content25 26# Introduction to Decision Tree27st.markdown("<h1 style='color: #003366;'>Decision Tree</h1>", unsafe_allow_html=True)28st.markdown("""29 A **Decision Tree** is a supervised machine learning algorithm used for both classification and regression tasks. 30 It builds a tree-like structure to make decisions based on feature values. The tree consists of:31 - **Root Node**: Represents the entire dataset.32 - **Internal Nodes**: Represent features that help in making decisions.33 - **Leaf Nodes**: Represent the final output (decision).34 35 It works like a series of **if-else** conditions to classify or predict values.36""", unsafe_allow_html=True)37 38# Entropy Explanation and Formula39st.markdown("<h2 style='color: #003366;'>Entropy: Measuring Randomness</h2>", unsafe_allow_html=True)40st.markdown("""41 **Entropy** is a measure of the randomness in the dataset. It helps us to measure the impurity or disorder within a set.42 43 The formula for entropy is:""")44st.image("entropy-formula-2.jpg", width=300)45 46 47st.markdown("""48 Where:49 - \( p(i) \) is the probability of each class in the data.50 51 Example:52 Let's consider a dataset with two classes: **Yes** and **No**, and their probabilities:53 - \( P(Yes) = 0.5 \)54 - \( P(No) = 0.5 \)55 56 The entropy will be calculated as:57 58 $$ H(Y) = - (0.5 \cdot \log_2(0.5) + 0.5 \cdot \log_2(0.5)) = 1 $$59 60 Maximum entropy is **1** for a binary classification problem when the data is perfectly balanced.61""", unsafe_allow_html=True)62 63# Gini Impurity Explanation and Formula64st.markdown("<h2 style='color: #003366;'>Gini Impurity: Measuring Impurity</h2>", unsafe_allow_html=True)65st.markdown("""66 **Gini Impurity** is another metric used to measure the impurity of the data. It ranges between 0 and 1, 67 where 0 means perfectly pure (only one class) and 1 means maximum impurity.68 69 The formula for Gini Impurity is:""")70 71st.image("gini.png",width=300)72st.markdown("""73 Where:74 - \( p(i) \) is the probability of each class in the dataset.75 76 Example:77 Let's take the same dataset as above:78 - \( P(Yes) = 0.5 \)79 - \( P(No) = 0.5 \)80 81 The Gini impurity is calculated as:82 83 $$ Gini(Y) = 1 - (0.5^2 + 0.5^2) = 0.5 $$84 85 A Gini impurity of **0.5** means that the dataset is impure, with equal distribution of classes.86""", unsafe_allow_html=True)87 88# Decision Tree Construction and Working89st.markdown("<h2 style='color: #003366;'>Decision Tree Construction</h2>", unsafe_allow_html=True)90st.markdown("""91 The **Decision Tree** is built based on the feature that minimizes the impurity measure (either Entropy or Gini Impurity). 92 The tree is constructed in a **top-down** fashion, and each decision (split) is based on a feature that provides the best information gain (reduction in impurity).93 94 For example, in a simple **binary classification**, a decision tree might evaluate the feature \( x_1 \) first. If \( x_1 \) is greater than a threshold, it goes to the left node; otherwise, it goes to the right.95 96 The process of splitting continues until either:97 - The node has pure data (impurity = 0).98 - There are no further splits to make (e.g., all data points belong to the same class).99""", unsafe_allow_html=True)100 101# Iris Decision Tree Image102st.markdown("<h2 style='color: #003366;'>Decision Tree for Iris Dataset</h2>", unsafe_allow_html=True)103st.markdown("""104 Below is a decision tree constructed for the **Iris Dataset** using the decision tree algorithm (CART). It has been built using the Gini Impurity or Entropy metric to classify the Iris flowers based on their features (such as petal length, petal width, etc.).105 106 Below is the visual representation of the decision tree based on the Iris dataset:107""", unsafe_allow_html=True)108 109 110st.image("dt1 (1).jpg", caption="Decision Tree for Iris Dataset", use_container_width =True)111 112# Decision Tree Training and Testing Phases for Classification113st.markdown("<h2 style='color: #003366;'>Training and Testing Phases for Classification</h2>", unsafe_allow_html=True)114st.markdown("""115 ### Training Phase (Classification):116 - The **training phase** for a decision tree classification involves learning from the dataset by iteratively choosing the best feature splits based on **Gini Impurity** or **Entropy**.117 - The tree is built starting from the root node, with each node representing a feature, and branches representing possible splits. 118 - The algorithm stops when the data at the node is pure (all instances belong to the same class) or other stopping criteria (e.g., max depth) are met.119 120 ### Testing Phase (Classification):121 - In the **testing phase**, the trained decision tree is used to classify new, unseen data.122 - Starting from the root, the test instance is passed down the tree following the conditions (splits) until it reaches a leaf node, where the predicted class label is given.123 124 **Example**: 125 For a new instance of the Iris dataset, the decision tree will classify it as one of the three species: Setosa, Versicolor, or Virginica based on its petal length, petal width, and other features.126""", unsafe_allow_html=True)127 128# Decision Tree Training and Testing Phases for Regression129st.markdown("<h2 style='color: #003366;'>Training and Testing Phases for Regression</h2>", unsafe_allow_html=True)130st.markdown("""131 ### Training Phase (Regression):132 - The **training phase** for a decision tree regression is similar to classification, but instead of calculating the class label, it calculates a continuous value (the predicted outcome).133 - The algorithm splits the data based on feature values to minimize the **Mean Squared Error (MSE)** between the true values and predicted values in the resulting branches.134 - The tree is constructed using recursive binary splits, and it stops when the node has sufficiently pure data (low MSE).135 136 ### Testing Phase (Regression):137 - During the **testing phase**, the decision tree model predicts continuous values for the test dataset. 138 - The value at each leaf node represents the mean of the values in that region of the feature space.139 140 **Example**:141 If we apply a decision tree regression model to predict the price of a house based on features like square footage, number of bedrooms, etc., the tree will predict the price for unseen data based on the average values of the training data in each leaf node.142""", unsafe_allow_html=True)143 144# Pre-Pruning Techniques145st.markdown("<h2 style='color: #003366;'>Pre-Pruning Techniques</h2>", unsafe_allow_html=True)146st.markdown("""147 **Pre-pruning** refers to the process of stopping the tree from growing before it reaches its full potential, 148 in order to prevent overfitting. The goal is to limit the tree's growth at an early stage.149 Some key pre-pruning techniques include:150 151 1. **Max Depth**: Set a limit on the depth of the tree. This can prevent the tree from growing too large and complex.152 2. **Min Samples Split**: This defines the minimum number of samples required to split an internal node. Increasing this value prevents the model from making overly specific splits.153 3. **Min Samples Leaf**: Defines the minimum number of samples required in a leaf node. This ensures that nodes have a sufficient number of samples before they are split.154 4. **Max Features**: Limits the number of features considered for splitting at each node. This can reduce overfitting by limiting the complexity of each split.155 156 Example:157 - **Max Depth**: Set the maximum depth to 3, which prevents the tree from splitting deeper and thus avoids overfitting.158 - **Min Samples Split**: Set this to 10, meaning that a node will only split if there are at least 10 samples in it.159""", unsafe_allow_html=True)160 161# Post-Pruning Techniques162st.markdown("<h2 style='color: #003366;'>Post-Pruning Techniques</h2>", unsafe_allow_html=True)163st.markdown("""164 **Post-pruning** refers to the process of creating the full decision tree first, and then removing branches that have little significance. This technique is used to reduce overfitting by removing parts of the tree that don't add value.165 166 Post-pruning involves:167 1. **Cost Complexity Pruning**: A technique where branches are removed if they do not improve the tree's performance.168 2. **Pruning based on Validation Data**: Post-pruning can also be based on a separate validation dataset to determine the optimal tree structure.169""", unsafe_allow_html=True)170 171# Feature Selection using Decision Tree172st.markdown("<h2 style='color: #003366;'>Feature Selection using Decision Tree</h2>", unsafe_allow_html=True)173st.markdown("""174 **Feature Selection** is the process of selecting the most important features to improve the model's performance and reduce overfitting. 175 In Decision Trees, **Feature Importance** is calculated based on the **impurity reduction** at each split.176 The formula for calculating feature importance is:""")177 178st.image("feature.png",width=500)179st.markdown("""180 Example:181 After building the decision tree, we can calculate the feature importance for each feature. 182 The higher the importance value, the more significant that feature is for making the predictions.183""", unsafe_allow_html=True)184 185st.markdown("<h2 style='color: #003366;'>I provided how to implement Decision Tree clearly in the below link</h2>",unsafe_allow_html=True)186st.markdown(187 "<a href='https://colab.research.google.com/drive/1SqZ5I5h7ivS6SJDwlOZQ-V4IAOg90RE7?usp=sharing' target='_blank' style='font-size: 16px; color: #003366;'>Open Jupyter Notebook</a>", 188 unsafe_allow_html=True189)190 191 