Team Ai
Apppublic

Indhu27/MachineLearning_Algorithms

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes
1KNN-Algorithm.py163 linesDownload Raw Back to pages
1import streamlit as st2st.set_page_config(page_title="KNN", page_icon="🤖", layout="wide")3st.markdown("""4    <style>5        .stApp {6            background-color: #4A90E2;7        }8        h1, h2, h3 {9            color: #003366; /* Adjust this if needed to match your background */10        }11        .custom-font, p {12            font-family: 'Arial', sans-serif;13            font-size: 18px;14            color: white; /* Making all inside text white */15            line-height: 1.6;16        }17    </style>18    """, unsafe_allow_html=True)19 20 21# Title of the Streamlit Application22st.markdown("<h1 style='color: #003366;'>K-Nearest Neighbors (KNN)</h1>",unsafe_allow_html=True)23 24# KNN Algorithm Theory25st.write("""26K-Nearest Neighbors (KNN) is a simple machine learning algorithm used for both **classification** and **regression** problems. It predicts the output based on the `K` nearest neighbors of a data point.27Key points:28- KNN is non-parametric Algorithm. 29- KNN doesn't learn anything at the time of training, it simply stores the training data and uses it to make predictions when a new data point comes in.30- KNN uses a distance metric (like Euclidean distance) to calculate how close points are to each other.31""")32 33st.markdown("<h2 style='color: #003366;'>Working of K-Nearest Neighbors (KNN)</h2>",unsafe_allow_html=True)34 35st.markdown("<h2 style='color: #003366;'>Training Phase (Same for Classification & Regression)</h2>",unsafe_allow_html=True)36st.write(37    """38    - KNN **does not learn** anything during training.39    - It **stores** the entire dataset and waits for a new data point.40    """41)42 43st.markdown("<h2 style='color: #003366;'>Testing Phase (Classification)</h2>",unsafe_allow_html=True)44st.write(45    """46    1. Choose the value of **K** (number of neighbors).47    2. Compute the distance between the new data point and all points in the training dataset.48       - Common distance metrics:49         - **Euclidean Distance** (Most common)50         - **Manhattan Distance**51         - **Minkowski Distance**52    3. Select the **K nearest neighbors**.53    4. Perform **majority voting**:54       - Count the occurrences of each class among the **K** neighbors.55       - Assign the most frequent class as the prediction.56    """57)58 59st.markdown("<h2 style='color: #003366;'>Testing Phase (Regression)</h2>",unsafe_allow_html=True)60st.write(61    """62    1. Choose the value of **K**.63    2. Compute the distance between the new data point and all training points.64    3. Select the **K nearest neighbors**.65    4. Compute the predicted value using:66       - **Mean (Average) of K neighbors' values** → Standard KNN Regression67       - **Weighted Mean (Closer neighbors have higher weights)** → Weighted KNN Regression68    """69)70 71# Overfitting, Underfitting, and Best Fit in KNN72st.markdown("<h2 style='color: #003366;'>Overfitting, Underfitting, and Best Fit in KNN</h2>",unsafe_allow_html=True)73st.write("""74- **Overfitting**: This happens when the model learns too much from the training data, including noise. It results in high accuracy on training data but poor performance on unseen data.75- **Underfitting**: This happens when the model is too simple, and it fails to capture the patterns in the data.76- **Best Fit**: This is when we find a balance between underfitting and overfitting. Choosing the right `k` value can help in achieving the best fit for the model.77""")78 79# Training Error and Cross Validation80st.markdown("<h2 style='color: #003366;'>Training Error and Cross Validation Error</h2>",unsafe_allow_html=True)81st.write("""82When choosing the best `k` value for KNN, we need to make sure our model is neither overfitting nor underfitting. Cross-validation helps in determining this:83- **Training Error**: The error we get when testing the model on training data.84- **Cross-validation (CV) Error**: The error when we test on a separate validation set (not the training data). This helps in selecting the best model.85- **High Training Accuracy** but **low CV Accuracy** means overfitting.86- **Low Training Accuracy** but **high CV Accuracy** means underfitting.87We want to find a `k` value that gives us low training and CV error.88""")89 90# Hyperparameter Tuning91st.markdown("<h2 style='color: #003366;'>Hyperparameter Tuning in KNN</h2>",unsafe_allow_html=True)92st.write("""93In KNN, the **main hyperparameter** is `k`, the number of neighbors to consider. Other hyperparameters include:94- **weights**: You can choose between `uniform` (each neighbor has equal weight) or `distance` (closer neighbors have higher weight).95- **distance metric**: You can use different distance metrics like **Euclidean** or **Manhattan** distance.96- **n_jobs**: This helps in using multiple CPU cores to speed up the training process.97Choosing the right value for `k` and other hyperparameters is important to get good performance.98""")99 100# Feature Scaling and Normalization101st.markdown("<h2 style='color: #003366;'>Feature Scaling</h2>",unsafe_allow_html=True)102st.write("""103Before using KNN, we need to make sure all features are on the same scale. This is important because KNN calculates distances between points, and features with larger scales will dominate the distance calculation. 104We can use:105- **Normalization** (Min-Max scaling): This scales data between 0 and 1.106- **Standardization** (Z-score scaling): This scales data to have a mean of 0 and a standard deviation of 1.107  108**Important**: Always apply scaling separately on training, validation, and test data to avoid data leakage.109""")110 111# Weighted KNN112st.markdown("<h2 style='color: #003366;'>Weighted KNN</h2>",unsafe_allow_html=True)113st.write("""114When data points are very close to each other, KNN might not perform well. **Weighted KNN** addresses this by giving closer neighbors more weight during classification.115For example:116- If `k=3`, and you have 3 neighbors with 2 being of class "Red" and 1 of class "Blue", KNN would predict the class as "Red".117- In **Weighted KNN**, closer neighbors get higher weights. So, if the closest neighbor is "Red", it will have more influence on the prediction.118""")119 120# KNN and Decision Regions121st.markdown("<h2 style='color: #003366;'>KNN and Decision Regions</h2>",unsafe_allow_html=True)122st.write("""123The KNN algorithm creates **decision regions** to classify data points. For example:124- When `k=1`, the decision boundary is highly sensitive, leading to overfitting (small, jagged regions).125- When `k` is larger, the decision region becomes smoother, reducing overfitting but increasing the risk of underfitting.126The key is to choose an optimal `k` that balances both.127""")128 129# Cross-Validation Explained130st.markdown("<h2 style='color: #003366;'>Cross Validation (CV)</h2>",unsafe_allow_html=True)131st.write("""132**Cross-Validation** is used to evaluate the model’s performance and prevent overfitting. It involves splitting the data into multiple parts (folds) and training the model on some parts while testing it on others. This helps in ensuring the model generalizes well on unseen data.133For example:134- **K-Fold Cross-Validation**: Split data into `k` folds and use each fold as a test set once. This helps in assessing the model performance more reliably.135""")136 137# Grid Search, Random Search, and Bayesian Search138st.markdown("<h2 style='color: #003366;'>Hyperparameter Tuning Techniques</h2>",unsafe_allow_html=True)139st.write("""140### Grid Search:141Grid Search tries every possible combination of hyperparameters to find the best ones. For example, if `k` can be {1, 2, 3, 4, 5} and weights can be {uniform, distance}, Grid Search will test all combinations (1, uniform), (1, distance), (2, uniform), and so on. It is exhaustive but can be slow for large datasets.142### Random Search:143Random Search randomly selects hyperparameters from the search space. It's faster than Grid Search and can still find good hyperparameters. The downside is that it might miss the best combination.144### Bayesian Search:145Bayesian Search uses a probabilistic model to select the next hyperparameters based on previous tests. It learns from past results to guide the search, making it more efficient than Grid and Random Search. This is especially useful when there are many hyperparameters to tune.146""")147 148st.markdown("<h2 style='color: #003366;'>I provided how to implement KNN clearly in the below link</h2>",unsafe_allow_html=True)149st.markdown(150    "<a href='https://colab.research.google.com/drive/11wk6wt7sZImXhTqzYrre3ic4oj3KFC4M?usp=sharing' target='_blank' style='font-size: 16px; color: #003366;'>Open Jupyter Notebook</a>", 151    unsafe_allow_html=True152)153    154 155 156 157 158 159st.write("""160KNN is a simple yet powerful algorithm, but it’s important to select the right hyperparameters, scale your features, and tune the model to avoid overfitting or underfitting. Cross-validation and hyperparameter tuning are key to getting the best performance out of the KNN model.161""")162 163