Indhu27/MachineLearning_Algorithms
0
1import streamlit as st2from sklearn.model_selection import train_test_split3import pandas as pd4 5st.set_page_config(page_title="Basic steps before training", page_icon="🤖", layout="wide")6st.markdown("""7 <style>8 .stApp {9 background-color: #4A90E2;10 }11 h1, h2, h3 {12 color: #003366; /* Adjust this if needed to match your background */13 }14 .custom-font, p {15 font-family: 'Arial', sans-serif;16 font-size: 18px;17 color: white; /* Making all inside text white */18 line-height: 1.6;19 }20 </style>21 """, unsafe_allow_html=True)22 23 24 25 26st.markdown("<h1 style='color: #003366;'>Basic steps to perform before training any Machine Learning Model</h1>", unsafe_allow_html=True)27 28st.write("""29### **Step 1: Understanding the Data**30- All columns in the dataset are called **variables**.31- Variables are divided into:32 - **Feature Variables (Independent Variables)** - These are the input variables (X).33 - **Target Variable (Dependent Variable)** - This is the output variable (Y).34""")35 36st.write("""37### **Step 2: Splitting the Data into Features (X) and Target (Y)**38- We separate the independent variables (X) and the dependent variable (Y) from the dataset.39 40**Example:** If we have a dataset with columns ['Age', 'Salary', 'Purchased']:41- Feature variables (X) = ['Age', 'Salary']42- Target variable (Y) = ['Purchased']43""")44 45st.write("""46### **Step 3: Splitting Data into Training and Testing Sets**47- The dataset is divided into two parts:48 - **Training Set (D_train)**: Used to train the model.49 - **Testing Set (D_test)**: Used to evaluate the model.50- We use **random splitting** to ensure that the training data does not repeat in the test set.51 52#### **Splitting Guidelines:**53- Majority of data points go into **D_train** (e.g., 70% or 80%).54- Minority of data points go into **D_test** (e.g., 30% or 20%).55- Splitting should be **random** to ensure fair distribution.56 57### **Implementation in Python**58""")59 60uploaded_file = st.file_uploader("Upload your dataset (CSV format)", type=["csv"])61if uploaded_file:62 df = pd.read_csv(uploaded_file)63 st.write("### Preview of Uploaded Data")64 st.write(df.head())65 66 target_column = st.selectbox("Select the Target Variable (Dependent Variable):", df.columns)67 feature_columns = st.multiselect("Select Feature Variables (Independent Variables):", [col for col in df.columns if col != target_column])68 69 if target_column and feature_columns:70 X = df[feature_columns]71 y = df[target_column]72 73 test_size = st.slider("Select Test Data Percentage", min_value=0.1, max_value=0.5, value=0.3)74 75 X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=test_size, random_state=42)76 77 st.write("### Data Split Summary")78 st.write(f"Training Data: {X_train.shape[0]} samples")79 st.write(f"Testing Data: {X_test.shape[0]} samples")80 81 st.write("#### Training Data Preview:")82 st.write(X_train.head())83 st.write("#### Testing Data Preview:")84 st.write(X_test.head())85 86 st.success("Data successfully split into training and testing sets!")87 