These are the answers to every question in Problem Set 2, written out in full. Read it alongside the problem set rather than instead of it, and attempt each question before you look here.
Everything below is examinable in Quiz 1.
Divide the top and the bottom by the same number:
\[\frac{S_{xY}}{S_{xx}} = \frac{S_{xY}/(n-1)}{S_{xx}/(n-1)} = \frac{\widehat{\mathrm{Cov}}(x, Y)}{\widehat{\mathrm{Var}}(x)}.\]
The \(n-1\) cancels because it
appears in both. This is why R’s cov(x,y)/var(x) gives the
slope even though neither cov nor var is a sum
of squares:
c(ratio_of_sums = Sxy / Sxx, ratio_of_moments = cov(x, y) / var(x))
## ratio_of_sums ratio_of_moments
## 3.932409 3.932409
Among the cars in this study, each additional mile per hour of speed is associated with an increase of about 3.93 feet in the mean stopping distance. The units are feet per mph.
Two things to be careful about. It is a statement about the mean: individual cars at the same speed stop at different distances, and \(\sigma^2\) measures this variability. Also, since the cars were observed and not assigned to specific values of speeds, the number describes how the whole data vary together.
Literally, \(\hat\beta_0 = -17.58\) says that a car travelling at 0 mph requires a stopping distance of about \(-17.6\) feet: it comes to rest roughly eighteen feet before it starts braking. This does not make sense.
What has gone wrong is extrapolation. The speeds in the data run from 4 to 25 mph, and \(x = 0\) is far outside that range. The intercept is where the fitted line happens to meet the vertical axis, which is a property of the line, but has no physical meaning. A straight line can be a good description of the data over the speeds observed while being useless outside that range. Actually, we know the stopping distance grows roughly with the square of speed, so the true relationship is curved, and a straight line fitted over 4 to 25 mph is a local approximation.
This is the standard warning about \(\hat\beta_0\): it is interpretable as a mean response only when \(x = 0\) is in the range of the observations.
With \(k_i = (x_i - \bar x)/S_{xx}\), the \(k_i\) depend only on the \(x\) values, which are fixed by the design. They contain no \(Y\) and are not random. So \(\hat\beta_1 = \sum_i k_i Y_i\) is a linear combination of the responses with constant coefficients, and expectation and variance are straighforward:
\[E(\hat\beta_1) = \sum_i k_i E(Y_i), \qquad \mathrm{Var}(\hat\beta_1) = \sum_i k_i^2 \mathrm{Var}(Y_i) \;+\; 2\sum_{i<j} k_i k_j \mathrm{Cov}(Y_i, Y_j).\]
That is the whole reason Lecture 4 can establish unbiasedness and the variance formula in a few lines.
sum(e) prints 1.1e-14The identity is exact, but the computer arithmetic is not.
R stores numbers in binary floating point with about 16 significant digits. Each residual is computed with a rounding error in the 16th digit, and 50 of them accumulate to something of order \(10^{-14}\). Against data of order \(10^2\), a discrepancy of \(10^{-14}\) reads as \(0\).
The practical rule: do not test a numerical identity with
==. Instead, compare against a tolerance, which is what
all.equal does:
e <- residuals(lm(dist ~ speed, data = cars))
c(raw = sum(e), equal_to_zero = isTRUE(all.equal(sum(e), 0)))
## raw equal_to_zero
## 1.110223e-14 1.000000e+00
With \(Q(\beta_0, \beta_1) = \sum_{i=1}^n (Y_i - \beta_0 - \beta_1 x_i)^2\),
\[\frac{\partial Q}{\partial \beta_0} = -2\sum_i (Y_i - \beta_0 - \beta_1 x_i), \qquad \frac{\partial Q}{\partial \beta_1} = -2\sum_i x_i (Y_i - \beta_0 - \beta_1 x_i).\]
Setting both to zero at \((\hat\beta_0, \hat\beta_1)\), and writing \(e_i = Y_i - \hat\beta_0 - \hat\beta_1 x_i\), gives
\[\sum_i e_i = 0 \qquad \text{and} \qquad \sum_i x_i e_i = 0.\]
All four are algebra. None of them requires any assumption about how the data were generated.
Both identities in the first bullet come from setting both partial derivatives to zero.
The consequence worth remembering: these identities are not evidence that the model fits or is any good. They hold for a dataset with a strong linear relationship and equally for one with no relationship at all. What does require assumptions: that \(\hat\beta_1\) is unbiased, that its variance is \(\sigma^2/S_{xx}\), and that \(\mathrm{MSE}\) estimates \(\sigma^2\) without bias.
That is a threat to constant variance: the Gauss-Markov condition \(\mathrm{Var}(\varepsilon_i) = \sigma^2\) assumes the same \(\sigma^2\) for every observation. Here the residuals fan out as speed increases, which is what you would expect if fast stops are more variable than slow ones.
We can list the consequences as:
Since \(\hat\beta_0 = \bar Y - \hat\beta_1 \bar x\),
\[e_i = Y_i - \hat\beta_0 - \hat\beta_1 x_i = (Y_i - \bar Y) - \hat\beta_1 (x_i - \bar x).\]
Square and sum:
\[\mathrm{SSE} = \sum_i e_i^2 = \underbrace{\sum_i (Y_i - \bar Y)^2}_{S_{YY}} - 2\hat\beta_1 \underbrace{\sum_i (x_i - \bar x)(Y_i - \bar Y)}_{S_{xY}} + \hat\beta_1^2 \underbrace{\sum_i (x_i - \bar x)^2}_{S_{xx}}.\]
Now use \(\hat\beta_1 S_{xx} = S_{xY}\) on the last term, so \(\hat\beta_1^2 S_{xx} = \hat\beta_1 S_{xY}\):
\[\mathrm{SSE} = S_{YY} - 2\hat\beta_1 S_{xY} + \hat\beta_1 S_{xY} = \boxed{S_{YY} - \hat\beta_1 S_{xY}}.\]
Two parameters, \(\beta_0\) and \(\beta_1\), were estimated from the data before the residuals could be formed. The residuals then satisfy the two normal equations \(\sum e_i = 0\) and \(\sum x_i e_i = 0\), which are two linear constraints, so only \(n - 2\) of the \(n\) residuals are free to vary. Dividing by the number of free residuals rather than by \(n\) is precisely what makes the estimator unbiased:
\[E(\mathrm{MSE}) = E\!\left(\frac{\mathrm{SSE}}{n-2}\right) = \sigma^2.\]
Dividing by \(n\) would give an estimator that is systematically too small, and by \(n-1\) one that is still too small, though less so.
\(\hat\sigma \approx 15.4\) feet can be seen as the typical vertical distance from a point to the line: roughly how much two cars travelling at the same speed differ in the distance they take to stop.
One foot is \(0.3048\) metres, so every \(Y_i\) is multiplied by \(c = 0.3048\). Going through the formulas:
| Quantity | Effect | Why |
|---|---|---|
| \(\hat\beta_1\) | multiplied by \(c\) | \(S_{xY} \to c\,S_{xY}\), \(S_{xx}\) unchanged. Units become metres per mph |
| \(\hat\beta_0\) | multiplied by \(c\) | \(\bar Y \to c \bar Y\) and \(\hat\beta_1 \bar x \to c \hat\beta_1 \bar x\) |
| \(\hat\sigma\) | multiplied by \(c\) | every residual is multiplied by \(c\) |
| \(r\) | unchanged | \(S_{xY} \to cS_{xY}\) and \(\sqrt{S_{YY}} \to c\sqrt{S_{YY}}\), so \(c\) cancels |
fit_m <- lm(I(dist * 0.3048) ~ speed, data = cars)
c(slope_ratio = coef(fit_m)[2] / b1,
r_feet = cor(x, y), r_metres = cor(x, y * 0.3048))
## slope_ratio.speed r_feet r_metres
## 0.3048000 0.8068949 0.8068949
The correlation is scale free, which is exactly why it can be compared across problems measured in different units and the slope cannot.