15 Linear Model

Consider a linear model Y=β0+β1X1+⋯+βp−1Xp−1+ε,ε∼N(0,σ2).

By MLE (see this example): β→^=(XTX)−1XTY→.
Common hypothesis testing:H0:βj=0 vs. H1:βj≠0.

Theorem

Under H0, the test statistic S=1||e→||2/[(n−p)σ2]⋅β^jVar(β^j)∼tn−p.
(Check definition of t distribution). Here e→=Y→−Xβ→^ is the residual.

Facts:

  1. Under H0, β^jVar(β^j)∼N(0,1).
  2. ||e→||2σ2∼χn−p2. (see here)
  3. ||e→||2⊥⊥β→^.

Today we will show that these facts imply S∼tn−p.


Conditional density of continuous random variables:
X,Y are continuous RVs, with joint density fX,Y. Given measurable set A∈F, how to calculate P(X∈A|Y=y0)?
For δ≪1,P(X∈A|y0<Y<y0+δ)=P[(X∈A),(y0<Y<y0+δ)]P(y0<Y<y0+δ)=∫A∫y0y0+δfX,Y(x,y)dydx∫y0y0+δfY(y)dy≈∫AfX,Y(x,y0)δdxfY(y0)δ=∫AfX,Y(x,y0)fY(y0)dx

Conditional Density

fX|Y=y0(x)=fX,Y(x,y0)fY(y0).

fX|Y=y0(x) is well defined if fY(y0)>0.

Independence

X⊥⊥Y⇔fX|Y=y(x)=fX(x),∀x∈R,∀y∈R,s.t.fY(y)>0.⇔fX,Y(x,y)=fX(x)fY(y),∀x,y∈R.

Law of Total Probability

fX(x)=∫−∞+∞fX,Ydy=∫−∞+∞fX|Y=y(x)fY(y)dy.
Application 1

X,Y∼i.i.dN(0,1). Let R=XY. We want to calculatefR(r)=∫−∞+∞fR|Y=y(r)fY(y)dy.
By independence of X,Y,(R|Y=y)=d(Xy|Y=y)=dXy.
Since T(X)=Xy is invertible and differentiable,fXy(r)=fX(ry)|d(ry)dr|=|y|fX(ry). ThereforefR(r)=∫−∞+∞|y|12πe−(ry)2212πe−y22dy=12π∫−∞+∞ye−y22(1+r2)dy=1π(1+r2),thus R∼Cauchy.

Application 2

X1,Y1,⋯,Yk∼i.i.dN(0,1). R=XY12+⋯+Yk2k. (Recall: G=Y12+⋯+Yk2∼Gamma(k2,12).) Note that(R|G=y)=d(Xg/k|G=g)=dXg/k.ThusfXg/k=gkfX(rgk).By law of total probabilityfR(r)=∫0+∞fR|G=y(r)fG(y)dy=∫0+∞gkfX(rgk)fG(y)dy=12πk12k/21Γ(k/2)∫0+∞g12(k+1)−1e−g2(r2k+1)dg=Γ((k+1)/2)Γ(k/2)1πk(1+r2k)−k+12∼tk.
(Recall Γ(12)=π,Γ(α)=(α−1)Γ(α−1).)

What happens as k→∞?
Y12,⋯,Yk2∼i.i.dGamma(12,12), E(Yi2)=1.
By SLLN, Y12+⋯+Yk2k→a.s.1, so tk→N(0,1).