More about Regression

Residuals

Residuals are the errors between observed and predicted values:

Residual = Observed Popularity – Predicted Popularity

residualsspotify

Interpretation of Coefficient (Simple Linear Regression)

The slope $\beta_1$:

  • Tells us how much our target, popularity, changes (on average) for each additional minute of track duration.

  • If $\beta_1 < 0$, longer songs tend to be less popular.

  • If $\beta_1 > 0$, longer songs tend to be more popular.

Assumptions

When building a linear regression model, it is important to check its assumptions. We will go deeper into what the assumptions are in question 5. If the assumptions are satisfied, we can trust the results of inference. If they are not, the results lose validity. The parameter estimates will not follow the expected distributions, which means hypothesis tests may give misleading accept/reject decisions. In other words: if you’re giving a linear regression model information that doesn’t meet its assumptions, it will give you invalid information back.

Parts of these explanations have been adapted from Applied Statistics with R (Dalpiaz, book.stat420.org).

Other Important Terms

  • Slope tells us the direction/magnitude of the relationship (duration vs. popularity).

  • Residuals show the difference between actual popularity and predicted popularity.

  • R² tells us how much of the variation in popularity is explained by predictors.

  • p-value for the slope tests whether the relationship is statistically significant or could be due to chance.

  • We can expand the model by adding more features (loudness, danceability, energy, valence, etc.) for better predictions → this is called Multiple Linear Regression which we will explain in question 2.