These are the answers to the discussion questions in Problem Set 1, written out in full.
Read it alongside the problem set rather than instead of it. Everything here is examinable in Quiz 1, in the week 4 tutorial.
Two points that caused trouble in the room, before the questions themselves:
rnorm(n, mean, sd) takes the standard
deviation, not the variance. To sample from \(N(10, 4)\) you write
rnorm(n, 10, 2).after_stat(density) is not decoration.
A count histogram and a density curve live on different vertical scales,
so overlaying one on the other is meaningless unless the histogram is
rescaled to have total area 1.The simulation printed a mean of 10.0047 against \(\mu = 10\), and a standard deviation of
0.4522 against \(\sigma/\sqrt{n}
= 2/\sqrt{20} = 0.4472\).
With \(X_1,\dots,X_n\) independent and identically distributed, \(E(X_i) = \mu\) and \(\text{Var}(X_i) = \sigma^2\):
\[E(\bar X) = \frac{1}{n}\sum_{i=1}^n E(X_i) = \mu,\]
\[\text{Var}(\bar X) = \frac{1}{n^2}\,\text{Var}\!\left(\sum_i X_i\right) = \frac{1}{n^2}\sum_i \text{Var}(X_i) = \frac{\sigma^2}{n}.\]
The middle step is the one to remember. Pulling the variance through the sum requires independence: it is the only place that assumption is used, and it is the assumption that fails most often in real data. Much of what goes wrong later in this course, when errors are correlated, goes wrong at exactly this line.
Variances of independent quantities add; standard deviations do not. Summing \(n\) terms multiplies the variance by \(n\), and dividing by \(n\) divides it by \(n^2\), leaving \(\sigma^2/n\). Taking the square root gives \(\sigma/\sqrt{n}\).
The practical reading: to halve the standard error you need four times as much data. Precision is expensive.
The centre stays at 10. The spread halves, from \(2/\sqrt{20} = 0.447\) to \(2/\sqrt{80} = 0.224\), so the histogram becomes half as wide and twice as tall. The shape is unchanged: still exactly normal.
A sum of independent normal random variables is exactly normal, so no limit theorem is involved.
Keep the three facts separate, because Question 2 depends on the distinction:
| Fact | Requires |
|---|---|
| \(E(\bar X) = \mu\) | nothing beyond a finite mean |
| \(\text{Var}(\bar X) = \sigma^2/n\) | independence and finite variance |
| \(\bar X\) is normal | a normal population |
Only the third needs normality. The first two hold for any population with a finite variance, which is what makes Question 2 possible.
For \(X_i\) independent and identically distributed with mean \(\mu\) and finite variance \(\sigma^2\),
\[\frac{\bar X - \mu}{\sigma/\sqrt{n}} \ \xrightarrow{\ d\ }\ N(0,1).\]
Two clarifications worth making explicit:
At \(n = 1\) the histogram is the population itself: a decaying exponential with its mode at 0 and a long right tail. By \(n = 5\) it is unimodal with a much shorter tail. By \(n = 30\) it looks symmetric. The skewness does not disappear, it shrinks like \(1/\sqrt{n}\).
The largest sample quantiles sit above the line, so the right tail is still heavier than a normal distribution’s. A histogram of 5,000 values cannot show this, because the discrepancy lives in the last handful of points, where the bars are one pixel high.
This is the real argument for QQ plots, and it is the diagnostic you will use all term on regression residuals.
\(E(\bar X) = \mu = 1\) exactly, for every \(n\), by the derivation in Question 1, which never used normality. The centring is exact; only the shape is asymptotic. This is the payoff of separating the three facts in Question 1(c).
Let \(Y_k = \sum_{i=1}^k X_i^2\) with \(X_i\) independent \(N(0,1)\). Then:
\[E(Y_k) = k\,E(X_1^2) = k\,\text{Var}(X_1) = k,\]
\[\text{Var}(Y_k) = k\,\text{Var}(X_1^2) = k\big(E(X_1^4) - [E(X_1^2)]^2\big) = k(3 - 1) = 2k,\]
using \(E(X^4) = 3\) for a standard normal.
Now the key observation: \(Y_k\) is itself a sum of \(k\) independent, identically distributed terms, so the central limit theorem applies to it directly. That is the entire content of the question. Nothing new is required, only the recognition that a chi-square variable is a sum.
The skewness of \(\chi^2_k\) is \(\sqrt{8/k}\): 2.83 at \(k = 1\), 1.26 at \(k = 5\), and 0.4 at \(k = 50\).
\(Y_1 = X_1^2 \ge 0\), so
\[Z_1 = \frac{Y_1 - 1}{\sqrt{2}} \ \ge\ -\frac{1}{\sqrt{2}} = -0.707.\]
The histogram is truncated on the left: a hard wall, not a tail that fades. No shifting and scaling of a non-negative variable can produce a normal distribution, which is why \(k = 1\) looks hopeless no matter how many replicates you draw.
Slower. \(\text{Exp}(1)\) has skewness 2 and looked acceptable by \(n = 30\), while \(\chi^2_1\) has skewness 2.83 and needs \(k \approx 50\) for comparable quality. Same theorem, different starting skewness. This is the concrete reason “\(n \ge 30\)” cannot be a theorem.
The simulation printed a mean slope of 20.0054 against
\(\beta_1 = 20\), and a standard
deviation of 0.3734 against \(\sigma/\sqrt{S_{xx}} = 0.3698\), with \(S_{xx} = 182.8\).
Everything follows from one observation: \(\hat\beta_1\) is linear in the \(Y_i\).
\[\hat\beta_1 = \frac{\sum_i (x_i - \bar x)(Y_i - \bar Y)}{S_{xx}} = \sum_i k_i Y_i, \qquad k_i = \frac{x_i - \bar x}{S_{xx}},\]
where the \(x_i\) are fixed constants, not random. The weights satisfy \(\sum_i k_i = 0\) and \(\sum_i k_i x_i = 1\), so with \(Y_i = \beta_0 + \beta_1 x_i + \varepsilon_i\):
\[E(\hat\beta_1) = \sum_i k_i(\beta_0 + \beta_1 x_i) = \beta_0 \underbrace{\sum_i k_i}_{0} + \beta_1 \underbrace{\sum_i k_i x_i}_{1} = \beta_1,\]
\[\text{Var}(\hat\beta_1) = \sigma^2 \sum_i k_i^2 = \sigma^2\,\frac{\sum_i (x_i - \bar x)^2}{S_{xx}^2} = \frac{\sigma^2}{S_{xx}}.\]
If you remember one sentence from this tutorial, make it this one: \(\hat\beta_1\) is a linear combination of the \(Y_i\) with fixed weights. The three results below are its consequences.
Only \(\varepsilon\) is redrawn. \(Y\) is random because the errors are, so \(\hat\beta_1\), being a function of \(Y\), is random too. The \(x\) values are fixed by the design of the experiment and carry no randomness at all.
Being able to say which symbols in a formula are random and which are fixed is the single most useful habit in this course.
\(E(\hat\beta_1) = \beta_1\): the estimator is unbiased.
Note that the simulated value 20.0054 is not exactly 20,
and should not be. With 2,000 replicates the Monte Carlo error is
roughly \(0.37/\sqrt{2000} \approx
0.0083\), and the discrepancy sits comfortably inside that.
Since \(\text{Var}(\hat\beta_1) = \sigma^2/S_{xx}\), a larger \(S_{xx}\) gives a more precise slope. So spread the \(x\) values out: observations far from \(\bar x\) carry more information about the slope than observations near it.
There is a cost, and it returns in Week 10. Extreme \(x\) values are precisely the high-leverage points, and they are also where a wrong functional form does the most damage. Precision and robustness pull in opposite directions.
Normal errors make \(\hat\beta_1\) exactly normal, because a linear combination of independent normal variables is normal. This is the closure property from Question 1(c), used again.
Without normality, \(\hat\beta_1\) is still unbiased with the same variance, since neither result used normality. It is also approximately normal for large \(n\), by the central limit theorem, because it is a sum of independent terms.
So the inference is robust in large samples, but the exactness is not. That distinction is worth carrying: it is why the assumption matters most when \(n\) is small.