Nagendra18/Matplotlib-Seaborn
0
1import streamlit as st2import numpy as np3import pandas as pd4import matplotlib.pyplot as plt5import seaborn as sns6 7st.title(":blue[Understanding Histograms]")8st.markdown("<hr style='border: 2px solid black; width: 50%; margin-left: auto; margin-right: auto;'>", unsafe_allow_html=True)9 10st.write("""11A **histogram** is a graphical representation used to visualize the distribution of a single numerical continuous dataset. It allows you to observe how frequently different ranges (bins) of values occur in the data.12""")13 14st.write("### Key Components of a Histogram:")15 16st.write("- **Y-Axis**: Represents the count or frequency of elements within each bin (the height of the bars).")17st.write("- **X-Axis**: Displays the bins, which represent ranges of numerical values in the dataset.")18 19st.write("### Creating Histograms:")20 21st.write("1. **Matplotlib**: You can create a histogram using the `plt.hist()` function.")22st.code("plt.hist(data, bins='number_of_bins')")23 24st.write("2. **Seaborn**: Use `sb.histplot()` for a similar purpose.")25st.code("sb.histplot(data, bins='number_of_bins')")26 27st.write("### Understanding Bins:")28st.write("""29- Bins divide the entire range of values into smaller intervals.30- The **bin width** can be calculated as follows:31 - **Formula**: `bin width = (max value - min value) / number of bins`32""")33 34 35 36 37 38st.title(":blue[Generating a Histogram with Matplotlib]")39st.write("""40```python41x = np.random.randint(1, 20, size=(100,))42plt.hist(x, bins=10, color='blue', alpha=0.7) 43plt.title("Histogram of Random Integers")44plt.xlabel("Value")45plt.ylabel("Frequency")46plt.grid(axis='y')47st.pyplot(plt)48```49""")50 51x = np.random.randint(1, 20, size=(100,))52st.write("Random Data:", x)53 54st.write("### Create a Histogram")55st.write("""56We use `plt.hist()` to create a histogram from the generated data. This function automatically calculates the frequency of each integer in the dataset and displays it.57""")58 59 60plt.hist(x, bins=10, color='blue', alpha=0.7) 61plt.title("Histogram of Random Integers")62plt.xlabel("Value")63plt.ylabel("Frequency")64plt.grid(axis='y')65st.pyplot(plt)66 67 68 69 70 71st.title(":blue[Creating a Density Histogram]")72 73 74st.write("""75We use `plt.hist()` with the `density=True` parameter to create a density histogram. This parameter normalizes the histogram so that the area under the histogram sums to 1.76- `denicity(percent)`77""")78st.write("""79```python80# denicity(percent)81plt.hist(x,density=True)82```83""")84st.image("https://cdn-uploads.huggingface.co/production/uploads/66be1362737c4ed890949fa1/8FAS6RGO0GfzNJxhBF3Fi.png")85 86 87 88 89 90st.title(":blue[Adjusting Bar Width in a Histogram]")91 92st.write("""93We can adjust the width of the bars in the histogram by setting the `rwidth` parameter in the `plt.hist()` function. The `rwidth` parameter specifies the relative width of the bars, where `1.0` means the bars will touch each other and values less than `1.0` will create space between the bars.94""")95 96st.write("""97```python98# width value99plt.hist(x,rwidth=0.9)100```101""")102st.image("https://cdn-uploads.huggingface.co/production/uploads/66be1362737c4ed890949fa1/HKXFeTHwB-tYq6nYEvp8m.png")103 104 105st.title(":blue[Customizing X-Axis Ticks in a Histogram]")106 107st.write("""108We can customize the x-axis tick marks by using the `plt.xticks()` function. In this case, we will set specific tick values to represent the intervals more clearly.109""")110 111 112st.write("""113```python114# use ticks115plt.hist(x,rwidth=0.9)116plt.xticks([ 1. , 2.8, 4.6, 6.4, 8.2, 10. , 11.8, 13.6, 15.4, 17.2, 19. ])117plt.show()118```119""")120st.image("https://cdn-uploads.huggingface.co/production/uploads/66be1362737c4ed890949fa1/4yhCjOmIIddeIk_aFtzEO.png")121 122 123 124 125 126 127 128st.title(":blue[Creating a Histogram with Custom Bins]")129 130st.write("""131We can specify the number of bins in the histogram using the `bins` parameter. In this case, we will set the number of bins to 3.132""")133st.write("""134```python135# your own bins136plt.hist(x,rwidth=0.9,bins=3)137plt.xticks([ 1. , 2.8, 4.6, 6.4, 8.2, 10. , 11.8, 13.6, 15.4, 17.2, 19. ])138plt.show()139```140""")141st.image("https://cdn-uploads.huggingface.co/production/uploads/66be1362737c4ed890949fa1/Ne7NJIDmo44IF60BW6wrT.png")142 143 144 145 146 147st.title(":blue[Creating a Histogram with Specified Ranges]")148 149st.write("""150In this example, we will demonstrate how to create a histogram using specified ranges with the `plt.hist()` function in Matplotlib.151""")152 153st.write("""154We can define the bin ranges in the histogram using the `bins` parameter. In this case, we will specify the ranges as [1, 8, 13, 19].155""")156st.write("""157```python158# range 159plt.hist(x,rwidth=0.9,bins=[1,8,13,19])160plt.xticks([ 1. , 2.8, 4.6, 6.4, 8.2, 10. , 11.8, 13.6, 15.4, 17.2, 19. ])161plt.show()162```163""")164st.image("https://cdn-uploads.huggingface.co/production/uploads/66be1362737c4ed890949fa1/nVIlUplLNyW8A_B7ttKOs.png")165 166 167 168 169st.title(":blue[Creating a Step Histogram]")170 171st.write("""172In this example, we will demonstrate how to create a step histogram using the `plt.hist()` function in Matplotlib. A step histogram displays the counts in a way that emphasizes the transitions between bins.173""")174 175st.write("""176We can create a step histogram by specifying the `histtype` parameter as `"step"`. This will display the histogram as an overlay of bars, allowing for a clear view of the frequency distribution.177""")178st.write("""179```python180# histtype(bar) 181# hist type is a step is called overlay of bar182plt.hist(x,rwidth=0.9,bins=9,histtype="step")183plt.xticks([ 1. , 2.8, 4.6, 6.4, 8.2, 10. , 11.8, 13.6, 15.4, 17.2, 19. ])184plt.show()185```186""")187st.image("https://cdn-uploads.huggingface.co/production/uploads/66be1362737c4ed890949fa1/Xwrk4qIHH2-3fNCENFEoh.png")188 189 190 191 192 193 194st.title(":blue[Comparative Histogram for Weight Comparison]")195 196st.write("""197In this example, we will create a comparative histogram to visualize the weight distribution of two groups: men and women. The `plt.hist()` function allows us to plot multiple datasets on the same histogram for comparison.198""")199 200st.write("""201We will create a histogram with specified bins to compare the weight distributions of men and women. The `label` parameter is used to add legends for clarity.202""")203st.write("""204```python205menw=np.random.randint(20,100,size=(30,))206womenw=np.random.randint(20,100,size=(30,))207plt.hist([menw,womenw],bins=[20,40,60,80,100],rwidth=0.94,label=["menw","womenw"])208plt.xlabel("weight")209plt.ylabel("count")210plt.title("weight comparision")211plt.legend()212plt.show()213```214""")215st.image("https://cdn-uploads.huggingface.co/production/uploads/66be1362737c4ed890949fa1/s5z-2ILNBhDhleUMCFxO0.png")216st.write("""217```python218st.header(":red[Comparative Histogram for orientation='horizontal']")219menw=np.random.randint(20,100,size=(30,))220womenw=np.random.randint(20,100,size=(30,))221plt.hist([menw,womenw], color=["r","k"],bins=[20,40,60,80,100],rwidth=0.94,label=["menw","womenw"], orientation='horizontal')222plt.xlabel("weight")223plt.ylabel("count")224plt.title("weight comparision")225plt.legend()226plt.show()227```228""")229st.image("https://cdn-uploads.huggingface.co/production/uploads/66be1362737c4ed890949fa1/3X5y0QFl_rrs57tqBW1zB.png")230 231 232 233 234st.title(":blue[Seaborn Histogram Examples]")235 236 237st.write("### Generate Random Data")238x = np.random.randint(1, 20, size=(100,))239st.write("Generated Data:", x)240 241 242st.write("### Example 1: Basic Histogram")243st.write("""244In this example, we create a basic histogram using Seaborn's `histplot()` function.245The x-axis shows the numerical values, while the histogram displays the frequency of each value.246""")247st.write("""248```python249plt.figure(figsize=(8, 4))250sns.histplot(x)251plt.xticks([1., 2.8, 4.6, 6.4, 8.2, 10., 11.8, 13.6, 15.4, 17.2, 19.])252plt.title("Basic Histogram")253st.pyplot(plt)254```255""")256plt.figure(figsize=(8, 4))257sns.histplot(x)258plt.xticks([1., 2.8, 4.6, 6.4, 8.2, 10., 11.8, 13.6, 15.4, 17.2, 19.])259plt.title("Basic Histogram")260st.pyplot(plt)261 262 263st.write("### Example 2: Specifying Bins")264st.write("""265Here, we specify the number of bins for the histogram. 266The `bins` parameter controls how many intervals the data is divided into.267""")268st.write("""269```python270plt.figure(figsize=(8, 4))271sns.histplot(x, bins=4)272plt.xticks([1., 2.8, 4.6, 6.4, 8.2, 10., 11.8, 13.6, 15.4, 17.2, 19.])273plt.title("Histogram with 4 Bins")274st.pyplot(plt)275```276""")277plt.figure(figsize=(8, 4))278sns.histplot(x, bins=4)279plt.xticks([1., 2.8, 4.6, 6.4, 8.2, 10., 11.8, 13.6, 15.4, 17.2, 19.])280plt.title("Histogram with 4 Bins")281st.pyplot(plt)282 283 284st.write("### Example 3: Specifying Bin Width")285st.write("""286In this example, we set the bin width to 6. 287The `binwidth` parameter defines how wide each bin will be, affecting the level of detail in the histogram.288""")289st.write("""290```python291plt.figure(figsize=(8, 4))292sns.histplot(x, binwidth=6)293plt.xticks([1., 2.8, 4.6, 6.4, 8.2, 10., 11.8, 13.6, 15.4, 17.2, 19.])294plt.title("Histogram with Bin Width of 6")295st.pyplot(plt)296```297""")298plt.figure(figsize=(8, 4))299sns.histplot(x, binwidth=6)300plt.xticks([1., 2.8, 4.6, 6.4, 8.2, 10., 11.8, 13.6, 15.4, 17.2, 19.])301plt.title("Histogram with Bin Width of 6")302st.pyplot(plt)303 304 305st.write("### Example 4: Customizing Palette and Color")306st.write("""307This example demonstrates how to customize the histogram's color using the `palette` and `color` parameters.308The `palette` allows you to choose a predefined color palette, and the `color` parameter lets you set a specific color.309""")310st.write("""311```python312plt.figure(figsize=(8, 4))313sns.histplot(x, binwidth=3, palette="dark", color=(0, 0.1, 0.6))314plt.xticks([1., 2.8, 4.6, 6.4, 8.2, 10., 11.8, 13.6, 15.4, 17.2, 19.])315plt.title("Histogram with Custom Color and Bin Width of 3")316st.pyplot(plt)317```318""")319plt.figure(figsize=(8, 4))320sns.histplot(x, binwidth=3, palette="dark", color=(0, 0.1, 0.6))321plt.xticks([1., 2.8, 4.6, 6.4, 8.2, 10., 11.8, 13.6, 15.4, 17.2, 19.])322plt.title("Histogram with Custom Color and Bin Width of 3")323st.pyplot(plt)324 325 326 327 328st.header(":blue[Kernel Density Estimate (KDE) Plot]")329st.write(" ",divider=True)330 331st.markdown("""332A **Kernel Density Estimate (KDE) plot** is a statistical technique used to estimate the probability density function of a continuous random variable. It's a useful way to visualize the distribution of data points and can be considered a smoothed version of a histogram. Hereโs a comprehensive overview of KDE plots:333""")334 335# What is a KDE Plot?336st.subheader("What is a KDE Plot?")337st.markdown("""338- **Kernel Density Estimation**: KDE is a non-parametric way to estimate the probability density function of a random variable. Unlike histograms, which can be sensitive to the choice of bin width and can produce discontinuous distributions, KDE provides a continuous curve that represents the data distribution.339- **Univariate Continuous Data**: A KDE plot is particularly effective for visualizing the distribution of a single continuous variable (univariate). It shows how data points are distributed across different values.340""")341 342st.subheader("How Does a KDE Plot Work?")343st.markdown("""3441. **Kernel Function**: The KDE algorithm uses a kernel function (such as Gaussian, Epanechnikov, etc.) to create a smooth curve. The kernel function determines the shape of the distribution around each data point.3452. **Bandwidth Selection**: The bandwidth is a crucial parameter in KDE, determining how smooth the resulting density estimate will be. A small bandwidth can lead to overfitting, capturing too much noise in the data, while a large bandwidth can oversmooth the data, obscuring important features.346 - **Rule of Thumb for Bandwidth**: Common methods for selecting bandwidth include Silverman's rule of thumb and Scott's method. These methods provide a starting point, but the choice may depend on the specific dataset and analysis goals.3473. **Probability Density Function**: The KDE generates a continuous probability density function (PDF), which represents the likelihood of different outcomes in the dataset. The area under the KDE curve sums to 1, reflecting the total probability.348""")349 350 351st.subheader("How to Interpret a KDE Plot")352st.markdown("""353- **Peaks**: The peaks in a KDE plot indicate areas where data points are concentrated. Higher peaks represent a higher density of data points.354- **Tails**: The tails of the plot provide insights into the distribution's behavior at the extremes. This can indicate the presence of outliers or the spread of data points.355- **Comparison with Histograms**: Unlike histograms, which can be jagged and dependent on bin selection, KDE plots provide a smoother representation of the data distribution.356""")357 358 359st.subheader("Creating a KDE Plot")360 361 362data = np.random.normal(loc=0, scale=1, size=1000) # Normal distribution363 364 365plt.figure(figsize=(10, 6))366sns.kdeplot(data, bw_adjust=0.5, fill=True, color='blue', alpha=0.5)367plt.title("Kernel Density Estimate (KDE) Plot")368plt.xlabel("Value")369plt.ylabel("Density")370plt.grid(True)371st.pyplot(plt) 372 373 374st.markdown("""375- **Data Generation**: In this example, we generate random data from a normal distribution using `np.random.normal()`.376- **Creating the KDE Plot**: The `sns.kdeplot()` function from the Seaborn library is used to create the KDE plot. 377 - `bw_adjust`: This parameter adjusts the bandwidth of the KDE. A value less than 1 results in a narrower KDE (more sensitive to fluctuations in data), while a value greater than 1 results in a wider KDE.378 - `fill`: This parameter, when set to `True`, fills the area under the KDE curve with color.379 - `color` and `alpha`: These parameters define the color and transparency of the fill.380""")381 382 383st.subheader("Use Cases of KDE Plots")384st.markdown("""385- **Data Exploration**: KDE plots help visualize the distribution of data, making them useful in exploratory data analysis.386- **Identifying Distribution Type**: By visualizing the data, you can assess whether it follows a specific distribution (e.g., normal, bimodal).387- **Comparing Distributions**: KDE plots can be overlaid to compare distributions of different datasets, allowing for easy visualization of similarities and differences.388""")389 390 391st.subheader("Limitations of KDE Plots")392st.markdown("""393- **Sensitivity to Bandwidth**: The choice of bandwidth can significantly affect the appearance of the KDE plot, leading to misinterpretations if not chosen carefully.394- **Not Suitable for All Data Types**: KDE is primarily for continuous data and may not be appropriate for categorical data or small sample sizes.395""")396st.markdown("""397KDE plots are powerful tools for visualizing the distribution of continuous univariate data. By providing a smooth estimate of the probability density function, they enable better understanding and exploration of data patterns, making them essential in statistical analysis and data visualization.398""")399 400 401 402 403st.title("Empirical Cumulative Distribution Function (ECDF) Plot")404 405 406st.markdown("""407An **Empirical Cumulative Distribution Function (ECDF) plot** is a statistical tool used to visualize the cumulative distribution of a dataset. Unlike a probability density function (PDF), which shows the likelihood of individual values, the ECDF displays the proportion of data points less than or equal to a given value. This makes it useful for understanding the distribution of data and comparing different datasets.408""")409 410 411st.subheader("What is an ECDF Plot?")412st.markdown("""413- **Empirical CDF**: The ECDF is a step function that represents the proportion of observations falling below each value in the dataset. It provides a non-parametric estimation of the cumulative distribution function.414- **Univariate Continuous Data**: An ECDF plot is particularly effective for visualizing the cumulative distribution of a single continuous variable (univariate). It illustrates how data points accumulate across different values.415""")416 417 418st.subheader("How Does an ECDF Plot Work?")419st.markdown("""4201. **Data Points**: The ECDF is constructed by sorting the data points in ascending order and plotting them against their corresponding cumulative proportions.4212. **Step Function**: The ECDF creates a step function where each step represents the addition of one data point. The height of each step corresponds to the cumulative proportion of data points.4223. **Range of Values**: The x-axis represents the values in the dataset, while the y-axis shows the cumulative proportion, ranging from 0 to 1.423""")424 425 426st.subheader("How to Interpret an ECDF Plot")427st.markdown("""428- **Cumulative Proportion**: The height of the ECDF at any given value indicates the proportion of data points that are less than or equal to that value.429- **Comparison of Datasets**: By overlaying ECDFs from different datasets, you can visually compare their distributions and identify differences in behavior.430- **Data Spread**: The steepness of the ECDF provides insights into the spread of the data. A steep slope indicates a concentration of values in that range, while a flat slope indicates fewer data points.431""")432 433 434st.subheader("Creating an ECDF Plot")435 436 437data = np.random.normal(loc=0, scale=1, size=1000) # Normal distribution438 439plt.figure(figsize=(10, 6))440sns.ecdfplot(data, color='blue')441plt.title("Empirical Cumulative Distribution Function (ECDF) Plot")442plt.xlabel("Value")443plt.ylabel("Cumulative Proportion")444plt.grid(True)445st.pyplot(plt) 446 447 448st.markdown("""449- **Data Generation**: In this example, we generate random data from a normal distribution using `np.random.normal()`.450- **Creating the ECDF Plot**: The `sns.ecdfplot()` function from the Seaborn library is used to create the ECDF plot.451""")452 453 454st.subheader("Use Cases of ECDF Plots")455st.markdown("""456- **Data Exploration**: ECDF plots help visualize the cumulative distribution of data, making them useful in exploratory data analysis.457- **Comparing Distributions**: ECDF plots can be overlaid to compare distributions of different datasets, allowing for easy visualization of similarities and differences.458- **Identifying Percentiles**: ECDF plots can help identify specific percentiles (e.g., median, quartiles) in the data distribution.459""")460 461 462st.subheader("Limitations of ECDF Plots")463st.markdown("""464- **Sensitivity to Sample Size**: For smaller sample sizes, the ECDF may not accurately reflect the underlying population distribution.465- **Less Informative for Bimodal Distributions**: In cases of multimodal distributions, the ECDF may not provide clear insights into the individual modes.466""")467 468st.markdown("""469ECDF plots are valuable tools for visualizing the cumulative distribution of continuous univariate data. By illustrating the proportion of data points below each value, they facilitate a better understanding of data patterns and enable comparisons between different datasets, making them essential in statistical analysis and data visualization.470""")471 472 473 474 475st.title("Distribution Plot (Combined Histogram and KDE)")476 477 478st.markdown("""479A **distribution plot** provides a comprehensive view of the distribution of a dataset by combining both a histogram and a kernel density estimate (KDE). This visualization helps in understanding the underlying frequency distribution and the smoothness of the data.480""")481 482st.subheader("What is a Distribution Plot?")483st.markdown("""484- **Histogram**: A histogram shows the frequency of data points within specified ranges (bins). It is useful for visualizing the shape and spread of the data.485- **Kernel Density Estimate (KDE)**: A KDE is a smooth curve that estimates the probability density function of the random variable. It provides a continuous representation of the distribution.486""")487 488 489st.subheader("How Does a Distribution Plot Work?")490st.markdown("""4911. **Data Points**: The distribution plot takes the dataset and creates a histogram with specified bins to visualize the frequency of values.4922. **Smoothness with KDE**: The KDE overlays the histogram, providing a smoother estimate of the distribution. It is computed using a kernel function and a bandwidth parameter.4933. **Combined View**: The resulting plot displays both the histogram and the KDE, allowing for an intuitive understanding of the data distribution.494""")495 496 497st.subheader("Creating a Distribution Plot")498 499 500data = np.random.normal(loc=0, scale=1, size=1000) # Normal distribution501 502 503plt.figure(figsize=(10, 6))504sns.histplot(data, bins=30, kde=True, color='blue', stat='density', alpha=0.5)505plt.title("Distribution Plot (Histogram + KDE)")506plt.xlabel("Value")507plt.ylabel("Density")508plt.grid(True)509st.pyplot(plt) 510 511 512st.markdown("""513- **Data Generation**: In this example, we generate random data from a normal distribution using `np.random.normal()`.514- **Creating the Distribution Plot**: The `sns.histplot()` function from the Seaborn library is used to create the histogram and KDE.515 - `bins`: This parameter specifies the number of bins for the histogram.516 - `kde`: When set to `True`, it overlays the kernel density estimate on the histogram.517 - `stat='density'`: This option normalizes the histogram to form a density plot.518 - `alpha`: This parameter sets the transparency of the histogram fill.519""")520 521 522st.subheader("Use Cases of Distribution Plots")523st.markdown("""524- **Data Exploration**: Distribution plots help visualize the shape and spread of data, making them useful in exploratory data analysis.525- **Identifying Distribution Type**: By visualizing the data, you can assess whether it follows a specific distribution (e.g., normal, bimodal).526- **Comparison of Datasets**: Distribution plots can be overlaid to compare distributions of different datasets, allowing for easy visualization of similarities and differences.527""")528 529 530st.subheader("Limitations of Distribution Plots")531st.markdown("""532- **Bin Sensitivity**: The choice of bin size can significantly affect the appearance of the histogram, potentially leading to misinterpretations.533- **KDE Sensitivity**: The bandwidth used for the KDE can also impact the smoothness of the estimate, potentially obscuring important features of the data.534""")535 536st.markdown("""537Distribution plots are powerful tools for visualizing the distribution of continuous univariate data. By combining histograms with kernel density estimates, they enable better understanding and exploration of data patterns, making them essential in statistical analysis and data visualization.538""")539 