← All projects

NBA Salary Prediction Model

"Scoring, minutes, defense and age explain 62% of the variation in NBA salaries, and 5-fold cross-validation holds that up"

Scatter plot of log salary against points per game with a fitted line
Log salary against points per game for the 184 players, with a least-squares line. Scoring alone explains about 55% of the variation.
Course: STAT 3220, Introduction to Regression Analysis (fall 2025)
Data: 184 players from the 2024–25 season averaging 20+ minutes a game (Basketball Reference stats and contracts, Statista franchise values)
Tools: R (lm, car, olsrr, caret): multiple linear regression on log salary, 5-fold cross-validation
Process: EDA, VIF checks for multicollinearity (dropped a PTS × FTA interaction at VIF ≈ 23), stepwise selection, nested F-tests for the categorical predictors, residual and influence diagnostics
Result: adjusted R² = 0.62; 5-fold cross-validation matched the in-sample fit (CV R² ≈ 0.62), so no sign of overfitting
Findings: points per game was by far the strongest predictor, followed by minutes, defensive stocks (steals + blocks) and age group (27–31 earned significantly more); awards and position added nothing once those were in the model

Why this dataset

A class project for STAT 3220, Introduction to Regression Analysis.

What the diagnostics caught

Cook's distance, leverage and studentized residuals flagged a handful of influential players: Tim Hardaway Jr., Chris Paul and Toumani Camara. All three were paid far below what their production predicts, around $2.3 million a year while playing 28 to 33 minutes a game. Rather than silently dropping them, we kept them in the final model and marked two for a follow-up sensitivity analysis. The diagnostics also pointed at what the model is missing: salary depends on things the box score doesn't record, like when a contract was signed, injuries, reputation and the league's salary-cap rules. That's where I'd take the model next, along with regularization (ridge or lasso) for the correlated predictors.

My contribution

I handled all of the coding and the data-science work for our team: preparing the dataset, and building, selecting, validating and diagnosing the regression models in R.

Residuals against fitted values for the final model
Residuals against fitted values for the final model. The labeled points are the players the model most overestimates.
Cook's distance for each player in the final model
Cook's distance for the final model: observations 161, 235 and 154 (Tim Hardaway Jr., Chris Paul and Toumani Camara) stand out.