TDM 10100: Project 7 — Fall 2026
-
Introduction to Linear Regression
Motivation: Linear regression is a simple but powerful method for predicting a quantitative (continuous) target variable.
Context: Even though newer models exist, linear regression remains popular because:
-
Results can be interpreted directly (coefficients show how predictors affect the target).
-
You can use a single predictor variable or multiple predictors.
-
Math behind the model is more transparent compared to black-box models.
-
You can model the relationship between a numeric response and one (or more) predictors.
Scope: We will use the Spotify Dataset (1986-2023) to demonstrate since our target variable (popularity), what we are trying to predict, is continuous (0-100).
Dataset(s)
This project will use the Spotify dataset:
-
/anvil/projects/tdm/data/spotify/linear_regression_popularity.csv
|
In this project, please only use 2 cores in your Jupyter Lab session: (do not use 4 cores or 16 cores for this project) |
Project Objectives
In this project, you will build a linear regression model to predict the popularity of a song using features such as duration, loudness, and many more features. We will start to learn about model assumptions. We will build very simple models and interpret the results.
You can read more about the Spotify data set here:
We have written an overview about Regression for you to read, before starting the project:
and some additional details here:
Questions
Question 1
Read in the data and print the first five rows of the dataset. Save the dataframe as myDF.
import pandas as pd
myDF = pd.read_csv("/anvil/projects/tdm/data/spotify/linear_regression_popularity.csv")
Use the code provided below, to drop the columns listed from myDF. After dropping them, print the columns still in the data.
Note: For more information on the drop function in pandas you can go here here.
drop_cols = [
"Unnamed: 0", "Unnamed: 0.1", "track_id", "track_name", "available_markets", "href",
"album_id", "album_name", "album_release_date", "album_type",
"artists_names", "artists_ids", "principal_artist_id",
"principal_artist_name", "artist_genres", "analysis_url", "duration_min"]
myDF = myDF.drop(columns=drop_cols)
Now we setup the variables to be used as your prediction target and features for the regression:
# Target and features
# Note: It is not a typing mistake that `y` is written in lowercase and `X` is written in uppercase. This is a common approach.
y = myDF["popularity"].copy()
X = myDF.drop(columns=["popularity"]).copy()
-
List the columns that remain in myDF, after removing the columns listed above, in
drop_cols. -
Print the shape of
XandyusingX.shapeandy.shape.
Question 2
Splitting the Data
Before building a model, we often partition the data into multiple subsets (training data and testing data), to serve distinct roles in the model development process. The most common partitioning scheme involves subsets:
-
Training data is what the model actually learns from. It’s used to find patterns and relationships between the features and the target.
-
Test data is completely held out until the very end. It gives us a final check to see how well the model is likely to perform on brand-new data it has never seen before.
Understanding the Subsets
In supervised learning, our dataset is split into predictors (X) and a target variable (y). We can further divide these into training, and test subsets to properly evaluate model performance and prevent overfitting.
|
For this project, we will only perform a single random train/test split using a fixed random seed. |
Now we create an 80/20 train/test split (use random_state=42).
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
It is worthwhile to check the shape and the head of each variable, to make sure that they were created correctly.
Print the shape and the head of each of these 4 variables: X_train, X_test, y_train, and y_test.
Now we generate a histogram of y_train (popularity) using the code provided below. Be sure to include helpful axis labels and a title for the plot.**
Note: See documentation on using .histplot in seaborn library here.
import matplotlib.pyplot as plt
import seaborn as sns
plt.figure(figsize=(8,5))
sns.histplot(y_train, bins=30, kde=True, color="skyblue")
plt.xlabel("_____") # Fill in a label for the x-axis
plt.ylabel("______") # Fill in a label for the y-axis
plt.title("_____") # Fill in a title for the histogram
plt.show()
-
Print the
shapeand theheadof each of these 4 variables:X_train,X_test,y_train, andy_test. -
Generate a histogram of
y_train(popularity).
Question 3
Examine the histogram in Question 2, and determine whether the distribution appears roughly symmetric. In 2 to 3 sentences, note your observations about its skewness and distribution. For instance, you might make a note about the mean and min and max values of y_train.
It might be helpful to explore the output of:
y_train.describe()
-
In 2 to 3 sentences, note your observations about the skewness and distribution of the histogram from Question 2.
Question 4
Using the code given below, generate a scatterplot of popularity versus duration (in minutes), and include a fitted regression line.
import matplotlib.pyplot as plt
import seaborn as sns
# Convert duration to minutes for training data
duration_min_train = X_train["duration_ms"] / 60000
plt.figure(figsize=(8,5))
sns.scatterplot(x=duration_min_train, y=y_train, alpha=0.6)
sns.regplot(x=duration_min_train, y=y_train, scatter=False, color="red", ci=None)
plt.xlabel("_____") # Fill in a label for the x-axis
plt.ylabel("______") # Fill in a label for the y-axis
plt.title("_____") # Fill in a title for the scatterplot
plt.show()
-
Generate a scatterplot of popularity versus duration (in minutes), and include a fitted regression line.
Question 5
In 2 to 3 sentences, for the scatterplot from Question 4, describe:
(1) the relationship you observe between the two variables, and
(2) how the regression line is constructed to represent the overall trend in the data. The residuals are the differences (i.e., the differences) between the actual observed values (the data itself) and the predicted values (from the regression line).
-
In 2 to 3 sentences, describe: (1) the relationship you observe between the two variables, and (2) how the regression line is constructed to represent the overall trend in the data (and the residuals are the differences from such a prediction).
A few closing notes:
Linear Regression Assumptions
Often we talk about the assumptions of this model, which are remembered by LINE.
-
Linear. The relationship between $Y$ and the predictors is linear.
-
Independent. The errors $\epsilon_i$ are independent.
-
Normal. The errors $\epsilon_i$ are normally distributed (the “error” around the line follows a normal distribution).
-
Equal Variance. The variance $\sigma^2$ of the errors is the same, across the prediction.
|
If you are a data science or statistics major, a solid understanding of these assumptions is frequently discussed in coursework and often asked about during interviews for data science roles! We encourage you to not only memorize these assumptions but also develop a clear understanding of their meaning and implications. |
References
Some explanations in this project have been adapted from other sources in statistics and machine learning, as listed below.
-
James, Gareth; Witten, Daniela; Hastie, Trevor; Tibshirani, Robert; Taylor, Jonathan. An Introduction to Statistical Learning: with Applications in Python. Springer Texts in Statistics, 2023.
-
Dalpiaz, David. Applied Statistics with R. Available at: book.stat420.org/
Submitting your Work
Once you have completed the questions, save your Jupyter notebook. You can then download the notebook and submit it to Gradescope.
-
firstname_lastname_project7.ipynb
|
You must double check your You will not receive full credit if your |