Team Ai
Modelpublic

BentoUniAcc/Stack_Overflow_Salary_Predicting_Model

sourceHugging Faceupdated 5mo agoView on Hugging Face
0likes
README.md383 linesDownload Raw Back to root
1---2language:3  - en4 5 6tags:7  - salary-prediction8  - regression9  - classification10  - clustering11  - tabular12  - scikit-learn13  - stack-overflow14  - developer-survey15  - feature-engineering16  - gradient-boosting17 18 19datasets:20  - stack-overflow-developer-survey-202521 22base_model: None23---24 25# Assignment 2 – Developer Salary Prediction26### Stack Overflow Developer Survey 2025 | Classification, Regression & Clustering27 28---29 30## Video31 32# <video src="https://huggingface.co/BentoUniAcc/Stack_Overflow_Salary_Predicting_Model/resolve/main/Data%20Analysis%20Assignment%202%20Video%20Project%201.mp4" controls="controls" style="max-width: 720px;"></video>33 34---35 36## Overview37 38This project uses the **Stack Overflow Annual Developer Survey 2025** (49,123 responses, 170 features) to predict a software developer's annual salary. The pipeline covers end-to-end data science: exploratory analysis, feature engineering, unsupervised clustering, regression, and multi-class classification.39 40**Research Question:** Can we predict a software developer's annual salary from their professional profile, and which factors matter most?41 42---43 44## Dataset45 46| Property | Value |47|----------|-------|48| Source | Stack Overflow Developer Survey 2025 (Kaggle) |49| Raw rows | 49,123 |50| Raw columns | 170 |51| Target column | `ConvertedCompYearly` (annual salary in USD) |52| Final feature count | 253 (after engineering + cluster feature) |53 54---55 56## Part 1 – Setup57 58- Environment: Google Colab compatible59- Reproducibility seed: `SEED = 42`60- Key libraries: `pandas`, `numpy`, `scikit-learn`, `matplotlib`61 62---63 64## Part 2 – Exploratory Data Analysis65 66### 2.1 Data Cleaning67 68- Removed rows with no salary value69- Clipped extreme outliers at the 1st and 99th percentile (final median salary ≈ $75K)70- Removed 44 columns with >60% missing values (170 → 126 columns), protecting the top-15 salary correlates regardless of missingness71 72### 2.2 Missing Value Analysis73 74 75![01_Bar_chart_top-40_columns_by_missing](https://cdn-uploads.huggingface.co/production/uploads/69d8c774af594a45bf54cc48/Yidxz76808ecCkC9bDRFM.png)76 77Several columns exceed 60% missingness and are dropped. The protected essential columns are retained despite high missingness and imputed later.78 79### 2.3 Descriptive Statistics80 81| Statistic | Value |82|-----------|-------|83| Median salary | ~$75,000 |84| Distribution | Right-skewed |85| Outlier treatment | 1st–99th percentile clip |86 87### 2.4 Salary Distribution88 89 90![02_Distribution_of_Annual_Developer_Salary](https://cdn-uploads.huggingface.co/production/uploads/69d8c774af594a45bf54cc48/pL0L7cvLkouusXvuTrFHp.png)91 92The raw distribution is heavily right-skewed with a long tail above $200K.93 94### 2.5 Research Questions & Findings95 96**Q1: Does coding experience predict salary?**97 98 99![03_Does_Coding_Experience_Predict_Salary](https://cdn-uploads.huggingface.co/production/uploads/69d8c774af594a45bf54cc48/98VvwfqURZ5RcnpdbIIku.png)100 101Salary increases steeply through the first 15–20 years of experience then flattens. There is wide variance at every experience level, suggesting experience alone is not sufficient to predict salary.102 103**Q2: How does education level affect salary?**104 105 106![04_How_Does_Education_Level_Affect_Salary](https://cdn-uploads.huggingface.co/production/uploads/69d8c774af594a45bf54cc48/ZE5DovlexhruZZ4n6oQm8.png)107 108Median salary rises with education level, but the gap between a Bachelor's and Master's degree is smaller than expected. Professional degrees and doctoral holders show the highest median salaries.109 110**Q3: Which countries pay developers the most?**111 112 113![05_Which_Countries_Pay_Developers_the_Most](https://cdn-uploads.huggingface.co/production/uploads/69d8c774af594a45bf54cc48/Ly4_3GC1EKt7agGdNPTD1.png)114 115The US dominates with a median salary roughly 2–3× the global median. Israeli, Western European and Australian developers cluster in a second tier, while developers in Asia and South America earn less.116 117**Q4: Do remote workers earn more?**118 119 120![06_Do_Remote_Workers_Earn_More](https://cdn-uploads.huggingface.co/production/uploads/69d8c774af594a45bf54cc48/xExDwYtGmmheimXVMfnby.png)121 122Fully remote developers show a slight salary premium over hybrid and in-office roles. The difference is modest, suggesting remote work correlates with higher-paying companies rather than being a direct cause.123 124**Q5: How does salary vary across developer roles?**125 126 127![07_How_Does_Salary_Vary_Across_Developer_Roles](https://cdn-uploads.huggingface.co/production/uploads/69d8c774af594a45bf54cc48/4zMNkvffCW-SZ0t3JT6-C.png)128 129C-Suite and ML/Data Science roles have the widest salary ranges and highest medians. Full-stack and front-end developers cluster around the global median with less variance.130 131### 2.6 Final Feature Selection (~20 Columns)132 133| Category | Features |134|----------|----------|135| Target | `ConvertedCompYearly` |136| Numeric | `YearsCode`, `WorkExp`, `JobSat`, `JobSatPoints_11`, `JobSatPoints_4` |137| Demographics | `Age`, `Country`, `EdLevel`, `MainBranch` |138| Work profile | `Employment`, `RemoteWork`, `DevType`, `OrgSize` |139| Tech & AI | `LanguageHaveWorkedWith`, `AISelect` |140| Learning | `LearnCodeChoose`, `SOVisitFreq` |141 142### 2.7 EDA Takeaways143 1441. **Salary** is right-skewed; median ~$75K after cleaning1452. **Work experience** (`WorkExp`) and **coding experience** (`YearsCode`) are the strongest numeric predictors1463. **Country** is the dominant signal — geography explains more variance than any other feature1474. **Remote work** carries a small positive premium1485. **Developer role and education** have meaningful but secondary effects149 150---151 152## Part 3 – Baseline Model153 154A simple Linear Regression trained on raw numeric columns only — no encoding, no feature engineering.155 156### Train/Test Split157- 80/20 random split, `SEED=42`158 159### Results160 161| Metric | Baseline |162|--------|----------|163| MAE    | $45,810  |164| RMSE   | $61,947  |165| R²     | 0.1598   |166 167### Predicted vs. Actual168 169 170![08_Plot_1_Predicted_vs_Actual](https://cdn-uploads.huggingface.co/production/uploads/69d8c774af594a45bf54cc48/6rJTbIavUrv0mQz2cZz4B.png)171 172The baseline model struggles with high earners — predictions cluster around the mean and fail to capture the upper salary range. The scatter is wide, consistent with an R² of only 0.16.173 174---175 176## Part 4 – Feature Engineering & Clustering177 178### Engineering Steps179 180| Step | Description |181|------|-------------|182| 4.1 Numeric features | Derived ratio/interaction features |183| 4.2 Ordinal encoding | `EdLevel`, `OrgSize` mapped to integers |184| 4.3 One-hot encoding | `Country`, `RemoteWork`, `Employment`, `MainBranch`, `AISelect`, `SOVisitFreq`, `Age`, `PrimaryDevType` |185| 4.4 Language flags | Binary flag for each of the top-10 programming languages |186| 4.5 Imputation & scaling | Median imputation + `StandardScaler` → 249 features |187 188### KMeans Elbow Method189 190 191![09_46_KMeans_Elbow](https://cdn-uploads.huggingface.co/production/uploads/69d8c774af594a45bf54cc48/BE9Tv-2Wu7kJuA1j7JNDd.png)192 193The inertia curve decreases gradually without a sharp elbow, reflecting the high-dimensional and overlapping nature of the data. k=4 was selected as a reasonable balance between cluster granularity and interpretability.194 195### Silhouette Score Comparison196 197 198![10_Silhouette_scores_across_k_for_each_clustering_method](https://cdn-uploads.huggingface.co/production/uploads/69d8c774af594a45bf54cc48/X4f8cT2a8e9baBeyVGmnA.png)199 200Silhouette scores are low across all values of k, confirming that natural cluster separation is weak in this dataset. Agglomerative clustering consistently outperforms KMeans, peaking around k=4.201 202### Three Clustering Algorithms203 204| Algorithm | k / params | Silhouette |205|-----------|-----------|------------|206| KMeans | k=4 | 0.0109 |207| DBSCAN | eps=5, min_samples=25 | 0.0912 (7 clusters) |208| Agglomerative Ward | k=4 | **0.0224** |209 210The data's high dimensionality (249 features) makes density-based clustering (DBSCAN) impractical — inter-point distances are too large for meaningful core-point detection.211 212### Cluster Visualisations (PCA 2D)213 214 215![11_48_Separate_scatter_plots](https://cdn-uploads.huggingface.co/production/uploads/69d8c774af594a45bf54cc48/DFPqTW9luodTCEWaLOJI1.png)216 217KMeans splits the data into four roughly equal blobs with significant overlap in the PCA projection. The clusters correspond loosely to salary level but boundaries are indistinct.218 219![12_48_Separate_scatter_plots](https://cdn-uploads.huggingface.co/production/uploads/69d8c774af594a45bf54cc48/GmVq3G3FJO2ucpXWGywAl.png)220 221DBSCAN classifies the vast majority of points as noise, forming 7 clusters. High dimensionality makes distance-based density estimation ineffective on this dataset.222 223 224![13_48_Separate_scatter_plots](https://cdn-uploads.huggingface.co/production/uploads/69d8c774af594a45bf54cc48/j6b9SMso4DovtAbKgqnSj.png)225 226Agglomerative clustering produces the clearest separation, isolating a distinct high-salary cluster on the right of the PCA plot. The four tiers align visually with low, mainstream, high-mid, and elite salary groups.227 228### Cluster Profiles – Agglomerative (Chosen)229 230| Cluster | Mean Salary | Median Salary | Count |231|---------|------------|---------------|-------|232| 0 | $86,041 | $74,000 | 18,192 |233| 1 | $109,980 | $95,000 | 4,069 |234| 2 | $27,574 | $13,949 | 877 |235| 3 | $101,735 | $93,387 | 317 |236 237**Winner: Agglomerative Ward (k=4)** — highest silhouette score and four interpretable salary tiers (low-income, mainstream, high-mid, elite).238 239### Cluster Feature Added240 241`cluster_id` one-hot encoded and appended → **253 final features**242 243---244 245## Part 5 – Improved Regression Models246 247Three models trained on the full 253-feature matrix (249 engineered features + 4 cluster dummies).248 249### Results250 251| Model | MAE | RMSE | R² |252|-------|-----|------|----|253| Baseline Linear Regression | $45,810 | $61,947 | 0.1598 |254| Improved Linear Regression | $30,688 | $44,314 | 0.5701 |255| Random Forest (200 trees) | $31,998 | $45,784 | 0.5411 |256| **HistGradientBoosting (300 iters)** | **$28,991** | **$43,039** | **0.5944** |257 258### Model Performance Comparison259 260 261![14_Comparison_table](https://cdn-uploads.huggingface.co/production/uploads/69d8c774af594a45bf54cc48/UBAzzh-I7z_S46Oq2vySN.png)262 263HistGradientBoosting wins on all three metrics. The jump from baseline to improved linear regression is dramatic — encoding Country alone accounts for the majority of the R² improvement from 0.16 to 0.57.264 265### Feature Importance266 267 268![15_Feature_importance_for_all_three_models](https://cdn-uploads.huggingface.co/production/uploads/69d8c774af594a45bf54cc48/mqxEApvV_TWicLuvpR2sh.png)269 270Country dummies (especially US) dominate feature importance across all three models. Work experience and years of coding rank consistently high. The cluster feature appears in the top 20 for linear regression, validating the clustering step.271 272### Winning Model – Predicted vs. Actual273 274 275![16_Declare_the_winner_based_on_R²_highest_and_MAE_lowest](https://cdn-uploads.huggingface.co/production/uploads/69d8c774af594a45bf54cc48/yNJ302EBP1qFzXOaT4awP.png)276 277The HistGradientBoosting model tracks the perfect-prediction diagonal much more closely than the baseline. It still under-predicts some very high earners above $300K but captures the mid-range salary distribution well.278 279### Discussion280 281- **Baseline → Improved Linear Regression (+0.41 R²):** One-hot encoding `Country` was the single biggest improvement. Geography is the dominant salary signal.282- **Random Forest vs. Linear Regression:** Non-linear feature interactions (e.g. senior developer × US location) are captured naturally by trees.283- **HistGradientBoosting wins:** Sequential boosting focuses on the hardest predictions. Natively handles missing values and is 10–100× faster than standard GradientBoosting.284- **Cluster feature:** Pre-computed salary-tier signal from Part 4 particularly boosts Linear Regression.285 286### Winner: HistGradientBoosting Regressor287 288| Metric | Value |289|--------|-------|290| MAE | $28,991 |291| RMSE | $43,039 |292| R² | 0.5944 |293 294---295 296## Part 6 – Winning Regression Model Export297 298The winning regression model is saved to `winning_model_regression.pkl`.299 300---301 302## Part 7 – Salary Classification Setup303 304The continuous salary target is binned into four ordered classes:305 306| Class | Label | Range (USD/year) |307|-------|-------|------------------|308| 0 | Low | < $30,000 |309| 1 | Mid | $30,000 – $90,000 |310| 2 | High | $90,000 – $160,000 |311| 3 | Very High | > $160,000 |312 313### Class Distribution314 315 316![17_Bar_chart](https://cdn-uploads.huggingface.co/production/uploads/69d8c774af594a45bf54cc48/lYp3HW7LWzddsmpPvnsDt.png)317 318Mid-salary developers make up nearly 40% of the dataset. Very High earners are the smallest class at 13.6%, creating a mild class imbalance that the models must handle.319 320### Salary Distribution per Class321 322 323![18_Salary_Distribution_per_Class](https://cdn-uploads.huggingface.co/production/uploads/69d8c774af594a45bf54cc48/DcYHbvYjHLwj9qqq07te8.png)324 325Each bin shows a clean salary range with minimal overlap at the boundaries, confirming the thresholds were well-chosen. The Very High class has the widest spread, reflecting high variability among top earners.326 327---328 329## Part 8 – Classification Models330 331Same 253-feature matrix as regression, with a stratified 80/20 train/test split.332 333### Precision vs. Recall & False Positives vs. False Negatives334 335**Recall is prioritised over precision** in this task. Misclassifying a developer into a lower salary tier (a false negative) carries real-world cost — under-negotiation, poor benchmarking, missed career leverage — whereas a false positive (over-predicting a tier) is relatively benign.336 337**False Negatives are more critical than False Positives.** Predicting "Mid" when a developer is truly "High" or "Very High" obscures their earning potential. For this reason, evaluation uses **weighted F1-score**, which balances precision and recall across all four classes with particular attention to recall in the minority tiers (Low and Very High).338 339 340### Results341 342| Model | Accuracy | F1 (weighted) |343|-------|----------|---------------|344| Logistic Regression | 0.598 | 0.597 |345| Random Forest (200 trees) | 0.605 | 0.597 |346| **HistGradientBoosting (300 iters)** | **0.611** | **0.610** |347 348### Classification Model Comparison349 350 351![19_Summary_table](https://cdn-uploads.huggingface.co/production/uploads/69d8c774af594a45bf54cc48/d8PLnQGx-aXyF4cHVlMXJ.png)352 353HistGradientBoosting leads on both accuracy and weighted F1, though the margin between all three models is narrow. The gap is larger on F1, reflecting better handling of the minority classes.354 355### Confusion Matrices356 357 358![20_graph](https://cdn-uploads.huggingface.co/production/uploads/69d8c774af594a45bf54cc48/C8kfKc3bb0tQLheqNEaBM.png)359 360All three models struggle most with the High class ($90K–$160K), frequently confusing it with Mid. HistGradientBoosting shows the best recall on the Low and Very High tiers — the most actionable classes — with misclassifications mostly occurring between adjacent salary bands.361 362### Per-Class Performance – HistGradientBoosting (Winner)363 364| Class | Precision | Recall | F1 |365|-------|-----------|--------|----|366| Low (<$30K) | 0.66 | 0.76 | 0.71 |367| Mid ($30K–$90K) | 0.69 | 0.59 | 0.64 |368| High ($90K–$160K) | 0.53 | 0.49 | 0.51 |369| Very High (>$160K) | 0.51 | 0.67 | 0.58 |370 371### Winner: HistGradientBoosting Classifier372 373Sequential boosting handles the sparse one-hot encoded feature space well, focuses capacity on the most difficult salary boundaries, and outperforms both Logistic Regression and Random Forest on accuracy and F1.374 375---376 377## Final Model Files378 379| File | Contents |380|------|----------|381| `winning_model_regression.pkl` | HistGradientBoosting Regressor (MAE $28,991, R² 0.59) |382| `winning_model_classifier.pkl` | HistGradientBoosting Classifier (Accuracy 0.61, F1 0.61) |383