SurajDataScientist/Exploratory_Data_Analysis_using_python_Libraries
0
1import streamlit as st2import pandas as pd3import numpy as np 4 5st.title("Pandas")6 7page = st.sidebar.radio("Choose topic", [8 "Introduction", 9 "Series", 10 "Data frame", 11 "read_csv", 12 "read_excel"13])14 15 16if page == "Introduction":17 st.title("Introduction")18 19 20 st.write("Pandas is one of the most important and widely used library of the python for data analysis, it is mainly used for data manipulation and exploratory data analysis.")21 st.image("https://miro.medium.com/v2/resize:fit:1080/1*xRcXl9YKgpnb0zvXLvZprg.jpeg", width = 305)22 23 st.write("In pandas we mainly work with datastructures (Series and DataFrames) and Tools which are used for Data analysis, Data cleaning and Data structuring")24 25 st.image("https://www.altexsoft.com/static/blog-post/2024/5/36514c36-b2f0-48a2-8072-c915c5f98dd5.webp",width = 700)26 27 st.write("Here it works with two kinds of data structures called series and Dataframe.")28 st.write("Series is a 1-dimensional labelled array and Data frame is a 2-dimensional labelled array, here what is common in these 2 datastructures are they are both labelled which makes the data more easier to identify and use it for analysis.")29 st.image("https://www.altexsoft.com/static/blog-post/2024/2/a2b6d6bd-898e-424f-98a8-50b3bdf775eb.png",width=600)30 31 32 33 34elif page == "Series": 35 st.header("Series")36 st.write("Series is a 1-dimensional labelled array. We should always imagine series as a single column vector.")37 st.markdown("""38 **Properties of Series:**39 1. It contains only **homogeneous** data 40 2. It is **mutable** 41 3. It is a **sequential** data type42 """)43 44 st.write("**Creating a series**")45 46 47 st.write("This will return you an empty series object.")48 with st.echo():49 import pandas as pd50 var1 = pd.Series([])51 st.code(var1, language="python")52 53 54 st.write("This will create basic simple series")55 with st.echo():56 var2 = pd.Series([1,2,3,4])57 st.code(var2, language='python')58 59 60 st.write(" Giving labels(which are customizable) index names to the values.")61 with st.echo():62 var3 = pd.Series([1,2,3,4],dtype=np.int8,index=['a','b','c','d'])63 st.code(var3, language='python')64 65 66 st.write("**Series attributes**")67 with st.echo():68 var2.ndim69 #to know the dimension of the series- It will always return 1.70 71 with st.echo():72 var2.shape73 # to know the shape of the series74 75 with st.echo(): 76 var2.size77 #to know the size of series78 79 with st.echo(): 80 c1 = var2.values81 # to convert series into an array82 st.code(c1, language='python')83 with st.echo():84 var2.dtype85 # to know the data type of series elements86 87 88 89 st.write("**Series methods**")90 with st.echo():91 d1 = pd.Series(range(100,400))92 st.code(d1, language='python')93 94 st.write(" head() - by default it will return the first 5 elemnts, to cutsomize number we should give number in parameter.")95 with st.echo():96 d1.head()97 st.code(d1.head(), language='python') 98 99 st.write("tail() - tail will return last 5 elements,by default it will return the last 5 elemnts, to cutsomize number we should give number in parameter.")100 with st.echo():101 d1.tail()102 st.code(d1.tail(), language='python')103 104 105 st.write("astype()- astype method will help us to return modify the data type of the series.")106 with st.echo():107 d1.astype(np.int16)108 st.code(d1.astype(np.int16), language='python')109 110 st.write("memory_usage() - It is used to check the how much memory is being utilized by the series.")111 with st.echo():112 d1.memory_usage()113 st.code(d1.memory_usage(), language='python')114 115 116 st.write("drop() - it is used to drop the certain values from series and here drop is done by accessing index of the series and mentioning index inside drop function parameter.")117 st.write("Let's look at this series d1 and compare after applying drop() method.")118 st.code(d1, language='python')119 with st.echo(): 120 d1.drop(labels=[2,4,6])121 st.code(d1.drop(labels=[2,4,6]), language='python')122 123 124 st.write("**Accessing the elements**")125 st.write("We can access the elements inside a series using loc[] and iloc[].")126 st.write("loc[]- loc is used while accessing elements by labels which are given by user/tempoarary labels.")127 st.write("iloc[] - iloc is used to access elements by permanent label/default label given by system.")128 129 130 with st.echo():131 data = pd.Series(range(10,15),index=["a","b","c","d","e"])132 st.code(data, language='python')133 134 st.write("iloc - can be used in 3 ways - 1.single element accessing, 2.Integer accessing, 3.Slicing")135 with st.echo():136 y1 = data.iloc[0]137 138 st.code(y1, language='python')139 with st.echo():140 y2 = data.iloc[[0,3,2,4]]141 142 st.code(y2, language='python') 143 144 with st.echo():145 y3 = data.iloc[1:4]146 147 st.code(y3, language='python') 148 149 150 st.write("loc - can also be used in 3 ways - 1.single element accessing, 2.Integer accessing, 3.Slicing")151 with st.echo():152 z1 = data.loc["a"]153 st.code(z1, language='python')154 155 with st.echo():156 z2 = data.loc[["a","d","c","b"]]157 158 st.code(z2, language='python') 159 160 with st.echo():161 z3 = data.loc["a":"d"]162 st.code(z3, language='python')163 164 165 