About this document

These are the answers to every question in Problem Set 2, written out in full. Read it alongside the problem set rather than instead of it, and attempt each question before you look here.

Everything below is examinable in Quiz 1.


Stage 1: the estimates

(a) Why the two formulas for \(\hat\beta_1\) agree

Divide the top and the bottom by the same number:

\[\frac{S_{xY}}{S_{xx}} = \frac{S_{xY}/(n-1)}{S_{xx}/(n-1)} = \frac{\widehat{\mathrm{Cov}}(x, Y)}{\widehat{\mathrm{Var}}(x)}.\]

The \(n-1\) cancels because it appears in both. This is why R’s cov(x,y)/var(x) gives the slope even though neither cov nor var is a sum of squares:

c(ratio_of_sums = Sxy / Sxx, ratio_of_moments = cov(x, y) / var(x))
##    ratio_of_sums ratio_of_moments 
##         3.932409         3.932409

(b) What the slope means

Among the cars in this study, each additional mile per hour of speed is associated with an increase of about 3.93 feet in the mean stopping distance. The units are feet per mph.

Two things to be careful about. It is a statement about the mean: individual cars at the same speed stop at different distances, and \(\sigma^2\) measures this variability. Also, since the cars were observed and not assigned to specific values of speeds, the number describes how the whole data vary together.

(c) Reading the intercept literally

Literally, \(\hat\beta_0 = -17.58\) says that a car travelling at 0 mph requires a stopping distance of about \(-17.6\) feet: it comes to rest roughly eighteen feet before it starts braking. This does not make sense.

What has gone wrong is extrapolation. The speeds in the data run from 4 to 25 mph, and \(x = 0\) is far outside that range. The intercept is where the fitted line happens to meet the vertical axis, which is a property of the line, but has no physical meaning. A straight line can be a good description of the data over the speeds observed while being useless outside that range. Actually, we know the stopping distance grows roughly with the square of speed, so the true relationship is curved, and a straight line fitted over 4 to 25 mph is a local approximation.

This is the standard warning about \(\hat\beta_0\): it is interpretable as a mean response only when \(x = 0\) is in the range of the observations.

(d) On the weights \(k_i\)

With \(k_i = (x_i - \bar x)/S_{xx}\), the \(k_i\) depend only on the \(x\) values, which are fixed by the design. They contain no \(Y\) and are not random. So \(\hat\beta_1 = \sum_i k_i Y_i\) is a linear combination of the responses with constant coefficients, and expectation and variance are straighforward:

\[E(\hat\beta_1) = \sum_i k_i E(Y_i), \qquad \mathrm{Var}(\hat\beta_1) = \sum_i k_i^2 \mathrm{Var}(Y_i) \;+\; 2\sum_{i<j} k_i k_j \mathrm{Cov}(Y_i, Y_j).\]

That is the whole reason Lecture 4 can establish unbiasedness and the variance formula in a few lines.


Stage 2: the fitted line and its residuals

(a) Why sum(e) prints 1.1e-14

The identity is exact, but the computer arithmetic is not.

R stores numbers in binary floating point with about 16 significant digits. Each residual is computed with a rounding error in the 16th digit, and 50 of them accumulate to something of order \(10^{-14}\). Against data of order \(10^2\), a discrepancy of \(10^{-14}\) reads as \(0\).

The practical rule: do not test a numerical identity with ==. Instead, compare against a tolerance, which is what all.equal does:

e <- residuals(lm(dist ~ speed, data = cars))
c(raw = sum(e), equal_to_zero = isTRUE(all.equal(sum(e), 0)))
##           raw equal_to_zero 
##  1.110223e-14  1.000000e+00

(b) The normal equations

With \(Q(\beta_0, \beta_1) = \sum_{i=1}^n (Y_i - \beta_0 - \beta_1 x_i)^2\),

\[\frac{\partial Q}{\partial \beta_0} = -2\sum_i (Y_i - \beta_0 - \beta_1 x_i), \qquad \frac{\partial Q}{\partial \beta_1} = -2\sum_i x_i (Y_i - \beta_0 - \beta_1 x_i).\]

Setting both to zero at \((\hat\beta_0, \hat\beta_1)\), and writing \(e_i = Y_i - \hat\beta_0 - \hat\beta_1 x_i\), gives

\[\sum_i e_i = 0 \qquad \text{and} \qquad \sum_i x_i e_i = 0.\]

(c) Which facts are algebra and which need assumptions

All four are algebra. None of them requires any assumption about how the data were generated.

  • \(\sum e_i = 0\) and \(\sum x_i e_i = 0\) are the normal equations themselves.
  • \(\sum \hat Y_i = \sum Y_i\) follows immediately from \(\sum e_i = 0\), since \(e_i = Y_i - \hat Y_i\).
  • The line passing through \((\bar x, \bar Y)\) is just \(\hat\beta_0 = \bar Y - \hat\beta_1 \bar x\) rearranged.

Both identities in the first bullet come from setting both partial derivatives to zero.

The consequence worth remembering: these identities are not evidence that the model fits or is any good. They hold for a dataset with a strong linear relationship and equally for one with no relationship at all. What does require assumptions: that \(\hat\beta_1\) is unbiased, that its variance is \(\sigma^2/S_{xx}\), and that \(\mathrm{MSE}\) estimates \(\sigma^2\) without bias.

(d) The widening spread

That is a threat to constant variance: the Gauss-Markov condition \(\mathrm{Var}(\varepsilon_i) = \sigma^2\) assumes the same \(\sigma^2\) for every observation. Here the residuals fan out as speed increases, which is what you would expect if fast stops are more variable than slow ones.

We can list the consequences as:

  • \(\hat\beta_1\) stays unbiased, because unbiasedness needs only \(E(\varepsilon_i) = 0\).
  • \(\mathrm{Var}(\hat\beta_1) = \sigma^2/S_{xx}\) is no longer correct, since it was derived by pulling out a single \(\sigma^2\) out of the sum.
  • So \(\mathrm{se}(\hat\beta_1)\) does not have the correct formula, and anything built on it, such as the confidence intervals and tests coming in week 3, is wrong with it.
  • Least squares is also no longer BLUE.

Stage 3: how far the points sit from the line

(a) Proving the shortcut

Since \(\hat\beta_0 = \bar Y - \hat\beta_1 \bar x\),

\[e_i = Y_i - \hat\beta_0 - \hat\beta_1 x_i = (Y_i - \bar Y) - \hat\beta_1 (x_i - \bar x).\]

Square and sum:

\[\mathrm{SSE} = \sum_i e_i^2 = \underbrace{\sum_i (Y_i - \bar Y)^2}_{S_{YY}} - 2\hat\beta_1 \underbrace{\sum_i (x_i - \bar x)(Y_i - \bar Y)}_{S_{xY}} + \hat\beta_1^2 \underbrace{\sum_i (x_i - \bar x)^2}_{S_{xx}}.\]

Now use \(\hat\beta_1 S_{xx} = S_{xY}\) on the last term, so \(\hat\beta_1^2 S_{xx} = \hat\beta_1 S_{xY}\):

\[\mathrm{SSE} = S_{YY} - 2\hat\beta_1 S_{xY} + \hat\beta_1 S_{xY} = \boxed{S_{YY} - \hat\beta_1 S_{xY}}.\]

(b) Why \(n - 2\)

Two parameters, \(\beta_0\) and \(\beta_1\), were estimated from the data before the residuals could be formed. The residuals then satisfy the two normal equations \(\sum e_i = 0\) and \(\sum x_i e_i = 0\), which are two linear constraints, so only \(n - 2\) of the \(n\) residuals are free to vary. Dividing by the number of free residuals rather than by \(n\) is precisely what makes the estimator unbiased:

\[E(\mathrm{MSE}) = E\!\left(\frac{\mathrm{SSE}}{n-2}\right) = \sigma^2.\]

Dividing by \(n\) would give an estimator that is systematically too small, and by \(n-1\) one that is still too small, though less so.

(c) What \(\hat\sigma\) describes

\(\hat\sigma \approx 15.4\) feet can be seen as the typical vertical distance from a point to the line: roughly how much two cars travelling at the same speed differ in the distance they take to stop.

(d) Changing the units of \(Y\)

One foot is \(0.3048\) metres, so every \(Y_i\) is multiplied by \(c = 0.3048\). Going through the formulas:

Quantity Effect Why
\(\hat\beta_1\) multiplied by \(c\) \(S_{xY} \to c\,S_{xY}\), \(S_{xx}\) unchanged. Units become metres per mph
\(\hat\beta_0\) multiplied by \(c\) \(\bar Y \to c \bar Y\) and \(\hat\beta_1 \bar x \to c \hat\beta_1 \bar x\)
\(\hat\sigma\) multiplied by \(c\) every residual is multiplied by \(c\)
\(r\) unchanged \(S_{xY} \to cS_{xY}\) and \(\sqrt{S_{YY}} \to c\sqrt{S_{YY}}\), so \(c\) cancels
fit_m <- lm(I(dist * 0.3048) ~ speed, data = cars)
c(slope_ratio = coef(fit_m)[2] / b1,
  r_feet = cor(x, y), r_metres = cor(x, y * 0.3048))
## slope_ratio.speed            r_feet          r_metres 
##         0.3048000         0.8068949         0.8068949

The correlation is scale free, which is exactly why it can be compared across problems measured in different units and the slope cannot.