Week 5
Data Analysis
Soci—316
Brief comments vis-à-vis your research memos are
available through Moodle.
Qualtrics Survey Assessment Deadline
Your Qualtrics survey assignments are due by
8:00 PM on Friday, Octobr 9th.
Qualtrics Survey Assessment
We should treat regressions as
counterfactual prediction machines.
Today, we’ll explore this idea—and work to demystify regressions—using a
basic example.
To predict the probability of surviving the sinking of the Titanic as a function of predictors like age, sex and
passenger class.
P(\text{Survival} = 1 \mid X) = \beta_0 + X\beta + \varepsilon
We’re going to estimate a
linear probability model.
Caveat Emptor
Using a linear probability model (LPM) to predict a binary outcome can produce predicted probabilities—P(Y=1\mid X)—that fall outside the [0, 1] interval.
We’re going to fit one anyway
(cf. Hippel 2015).
# A tibble: 9 × 3
name survived age
<chr> <dbl> <dbl>
1 Aubart, Mme. Leontine Pauline 1 24
2 Stone, Mrs. George Nelson (Mart 1 62
3 Cardeza, Mr. Thomas Drake Marti 1 36
4 Buss, Miss. Kate 1 36
5 Smith, Miss. Marion Elsie 1 40
6 Hocking, Mr. Richard George 0 23
7 Cohen, Mr. Gurshon Gus 1 18
8 Cribb, Mr. John Hatfield 0 44
9 Skoog, Miss. Mabel 0 9
P(\text{Survival} = 1 \mid X) = \beta_0 + \beta_1\text{age} + \varepsilon
We’re modelling the probability of survival as a function of age.
To fit this model, we need to calculate the slope and intercept parameters—that is, \beta_0 and \beta_1\text{age}, respectively.
Slopes can be derived using the following formula:
b = \frac{\sum{(X - \bar{X})(Y - \bar{Y})}}{\sum{(X - \bar{X})^2}}
Intercepts can be calculated as follows:
\alpha = \bar{Y} - b\bar{X}
| Y | X | X - \bar{X} | Y - \bar{Y} | (X - \bar{X})(Y - \bar{Y}) | (X - \bar{X})^2 |
|---|---|---|---|---|---|
| 1 | 24 | −8.44 | 0.33 | −2.81 | 71.31 |
| 1 | 62 | 29.56 | 0.33 | 9.85 | 873.53 |
| 1 | 36 | 3.56 | 0.33 | 1.19 | 12.64 |
| 1 | 36 | 3.56 | 0.33 | 1.19 | 12.64 |
| 1 | 40 | 7.56 | 0.33 | 2.52 | 57.09 |
| 0 | 23 | −9.44 | −0.67 | 6.30 | 89.20 |
| 1 | 18 | −14.44 | 0.33 | −4.81 | 208.64 |
| 0 | 44 | 11.56 | −0.67 | −7.70 | 133.53 |
| 0 | 9 | −23.44 | −0.67 | 15.63 | 549.64 |
| \sum (X - \bar{X}) = 0 | \sum (Y - \bar{Y}) = 0 | \sum (X - \bar{X})(Y - \bar{Y}) = 21.33 | \sum (X - \bar{X})^2 = 2008.22 |
\beta_1 = \frac{\sum (X - \bar{X})(Y - \bar{Y})}{\sum (X - \bar{X})^2}
{} = \frac{21.33}{2008.22} \approx 0.0106
\beta_0 = 0.667 - (0.0106 \times 32.44) \approx 0.3220
Our first LPM can therefore be summarized as follows:
\hat{P}(\text{Survived} = 1) = 0.3220 + 0.0106 \cdot \text{age} + \varepsilon
We can use programs like to streamline hand calculations.
library(tidyverse)
titanic_subset |> mutate(# Averages
mean_y = mean(survived),
mean_x = mean(age),
# To establish covariation between x and y
# Deviations from average, x
x_minus_mean_x = age - mean(age),
# Corresponding y value--deviations from average
y_minus_mean_y = survived - mean(mean_y),
# Tapping co-movement
slope_numerator = x_minus_mean_x * y_minus_mean_y,
# Squared residuals, positive values are yielded
slope_denominator = x_minus_mean_x^2) |>
# Putting it all together
mutate(# Summing up the values for the numerator ...
slope_numerator = sum(slope_numerator),
# And the denominator ...
slope_denominator = sum(slope_denominator),
# Yielding the slope estimate
slope = slope_numerator/slope_denominator,
# Which allows us to "solve" for the intercept:
intercept = mean_y - slope * mean_x) |>
select(intercept, slope) |>
distinct()# A tibble: 1 × 2
intercept slope
<dbl> <dbl>
1 0.322 0.0106
Kate Buss was 36 years old when the Titanic sank.
# A tibble: 1 × 3
name survived age
<chr> <dbl> <dbl>
1 Buss, Miss. Kate 1 36
We can compute her probability of survival Using our clearly misspecified model. as follows:
Or we can use functions to generate these predictions for us.
library(marginaleffects)
avg_predictions(survival_truncated, variables = "age") |>
as_tibble() |>
ggplot(mapping = aes(x = age, y = estimate)) +
geom_line(colour = "#b7a5d3", linewidth = 1.5) +
geom_ribbon(mapping = aes(ymin = conf.low,
ymax = conf.high),
fill = "#b7a5d3",
alpha = 0.35) +
theme_bw()survival_full <- lm(survived ~ age, data = titanic)
avg_predictions(survival_full, variables = "age") |>
as_tibble() |>
ggplot(mapping = aes(x = age, y = estimate)) +
geom_line(colour = "#b7a5d3", linewidth = 1.5) +
geom_ribbon(mapping = aes(ymin = conf.low,
ymax = conf.high),
fill = "#b7a5d3",
alpha = 0.35) +
theme_bw()Uh-oh.
Our initial model painted a misleading portrait.
Call:
lm(formula = survived ~ age + sex + passengerClass, data = titanic)
Residuals:
Min 1Q Median 3Q Max
-1.09442 -0.24408 -0.08375 0.23289 0.99397
Coefficients:
Estimate Std. Error t value Pr(>|t|)
(Intercept) 1.1049549 0.0438210 25.215 < 2e-16 ***
age -0.0052695 0.0009316 -5.656 2.00e-08 ***
sexmale -0.4914131 0.0255520 -19.232 < 2e-16 ***
passengerClass2nd -0.2113738 0.0348568 -6.064 1.85e-09 ***
passengerClass3rd -0.3703874 0.0325039 -11.395 < 2e-16 ***
---
Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
Residual standard error: 0.3913 on 1041 degrees of freedom
(263 observations deleted due to missingness)
Multiple R-squared: 0.3691, Adjusted R-squared: 0.3667
F-statistic: 152.3 on 4 and 1041 DF, p-value: < 2.2e-16
\begin{align*} \hat{P}(\text{Survived} = 1) = 1.105 - 0.00527 \cdot \text{age} - 0.4914 \cdot \text{sex}_\text{male} \\ - 0.2114 \cdot \text{class}_\text{2nd} - 0.3704 \cdot \text{class}_\text{3rd} + \varepsilon \end{align*}
Call:
lm(formula = survived ~ age + sex + passengerClass, data = titanic)
Coefficients:
(Intercept) age sexmale passengerClass2nd
1.104955 -0.005269 -0.491413 -0.211374
passengerClass3rd
-0.370387
We’re now modelling the probability of survival as a function of
age, sex and passenger class.
Here are three passengers selected at random:
# A tibble: 3 × 5
name survived sex age passengerClass
<chr> <dbl> <fct> <dbl> <fct>
1 Veal, Mr. James 0 male 40 2nd
2 Petersen, Mr. Marius 0 male 24 3rd
3 Johnson, Master. Harold Theodor 1 male 4 3rd
In small groups, calculate the probability of survival for each passenger.
Call:
lm(formula = survived ~ age + sex + passengerClass, data = titanic)
Coefficients:
(Intercept) age sexmale passengerClass2nd
1.104955 -0.005269 -0.491413 -0.211374
passengerClass3rd
-0.370387
# A tibble: 3 × 2
name estimate
<chr> <dbl>
1 Veal, Mr. James 0.191
2 Petersen, Mr. Marius 0.117
3 Johnson, Master. Harold Theodor 0.222
SE(\hat{\beta}) = \sqrt{\frac{\hat{\sigma}^2}{\sum_{i = 1}^n (x_i- \bar{x})^2}}
Parameter estimates based on our full model.
# A tibble: 1 × 5
term estimate std.error statistic p.value
<chr> <dbl> <dbl> <dbl> <dbl>
1 age -0.00527 0.000932 -5.66 0.0000000155
t = \frac{\hat{\beta} - \color{#856cb0}{\beta_{H_\theta}}}{SE(\hat{\beta})}
t = \frac{\hat{\beta} - \color{#856cb0}{\beta_{H_0}}}{SE(\hat{\beta})}
# A tibble: 1 × 6
term estimate std.error statistic manual_statistic p.value
<chr> <dbl> <dbl> <dbl> <dbl> <dbl>
1 age -0.00527 0.000932 -5.66 -5.66 0.0000000200
A result is statistically significant at the 5% level when the corresponding parameter is distinguishable from zero using 5% as the rejection threshold. Conversely, a result is not statistically significant at the 5% level when the corresponding parameter is not distinguishable from zero using 5% as the rejection threshold.
(Llaudet and Imai 2023:215, EMPHASIS ADDED)
As Llaudet and Imai (2023) note, we can use simple linear models to approximate the difference-in-means estimator and derive our key estimand, the \widehat{\textit{ATE}}.
In small groups, discuss how conceptualizing regressions as counterfactual prediction machines can help us understand what the \widehat{\textit{ATE}} represents.
For the rest of today’s session, please work on your
Qualtrics survey assignment.
Grimmer and colleagues (2021:396)
A class of flexible algorithmic and statistical techniques for prediction and dimension reduction.
Molina and Garip (2019:28)
A way to learn from data and estimate complex functions that discover representations of some input (X), or link the input to an output (Y) in order to make predictions on new data.
Once deployed, SML algorithms learn the complex patterns linking X—a set of features Independent variables or predictors. —to a target variable Outcome or dependent variable. , Y.
The goal of SML is to optimize predictions—i.e., to find functions or algorithms that offer substantial predictive power when confronted with
new or unseen data.
SML algorithms include, but are not limited to, logistic regressions, random forests, ridge regressions, support vector machines and neural networks.
A quick note on terminology:
If a target variable is quantitative, we are dealing with a regression problem.
If a target variable is qualitative, we are dealing with a classification problem.
# A tibble: 80,000 × 6
fruit shape weight_g colour texture origin
<chr> <chr> <int> <chr> <chr> <chr>
1 apple spherical 150 red crispy china
2 banana curved 120 yellow creamy ecuador
3 orange spherical 148 orange juicy egypt
4 watermelon spherical 4500 green juicy spain
5 strawberry conical 13 red juicy mexico
6 grape spherical 5 green juicy chile
7 mango ellipsoidal 240 yellow juicy india
8 pineapple conical 2100 yellow juicy costa rica
9 apple spherical 140 green crispy usa
10 banana curved 110 yellow creamy ecuador
# ℹ 79,990 more rows
# A tibble: 1 × 6
fruit shape weight_g colour texture origin
<chr> <chr> <int> <chr> <chr> <chr>
1 <NA> conical 6 red juicy china
strawberry71% probabilityUML algorithms search for a representation of the inputs/features that is more useful than X itself (Molina and Garip 2019).
With UML, analysts mine for hidden structure or latent patterns in high-dimensional space.
What does this mean, exactly?
With UML, there is no observed outcome variable Y—or target—to supervise the estimation process. Instead, we only have inputs.
The primary goal in UML is to develop a lower-dimensional representation of complex data by inductively learning from the
interrelationships among inputs.
Reducing a vector of features to a small set of scales (e.g., via principal component analysis) or dividing a sample into a small number of hidden groups (e.g., via k-means clustering).
V-Party Variable
Underlying question
v2paanteliv2papeoplev2paculsupv2paminorv2paplurv2paoprespv2paviolKarim and Lukk’s The Radicalization of Mainstream Parties in the 21st Century
As Grimmer and colleagues note (2021:397, EMPHASIS ADDED) note, “machine learning is as much a culture defined by a distinct set of values and tools as it is a set of algorithms.”
This point has, of course, been made elsewhere.
The terms generative and predictive (as opposed to data and algorithmic) come from David Donoho’s (2017) 50 Years of Data Science.
\hat{\beta} Generative Classical statistics
Inferring relationships between X and Y
\hat{y} Predictive Machine learning
Generating accurate predictions of Y
To be sure, the putative strengths and weaknesses of these modelling “cultures” have been hotly debated.
Advances in machine learning can provide empirical leverage to social scientists and sharpen social theory in one fell swoop.
Lundberg, Brand and Jeon (2022), for instance, argue that embracing machine learning can help social scientists in at least four ways.
ML is often associated with induction. van Loon (2022) argues that SML algorithms can help us deductively resolve predictability hypotheses, too.
Emerges when we build SML algorithms that fail to sufficiently map the patterns—or pick up the empirical signal—linking X and Y. Think: underfitting.
Arises when our algorithms not only pick up the signal linking X and Y, but some of the noise in our data as well. Think: overfitting.
To strike the optimal balance between bias and variance.
In an SML setting, we want to reduce our algorithm’s generalization or test error—i.e., “the prediction error of a model on new data” (Molina and Garip 2019).
To arrive at an estimate of our model’s performance, we can (randomly) partition our global sample of observations into disjoint sets or subsamples.
F1 score for classification problems)—or to derive our generalization error.Unlike standard train/test splits, v or k-fold cross-validation lets us learn from all our data.
k-fold cross-validation proceeds as follows:
Stratified k-fold cross-validation ensures that the distribution of class labels—or for numeric targets, the mean—is relatively constant across folds.
Figure 8 from Mittleman (2022).
Figure 8 from Boutyline and colleagues (2023)
Figure 2 from Karim (2026)
For the rest of today’s session, please work on your
Qualtrics survey assignment once again.
Note: Scroll to access the entire bibliography
