16 Bayesian Inference

?

Draw the region like in Lecture 16, we have P(U1+U2≤1)=12.
Draw a 3-d region, we have P(U1+U2+U3≤1)=16.
Consider the following polytope in Rn:Δn={(u1,⋯,un)∈Rn|0≤ui≤1,u1+⋯+un≤1}.Then Vol(Δn)=1n!. (We can use multi-dimensional integration & induction to show it.) So P(U1+⋯+Un≤1)=1n!.
Define Ek as the event of U1+⋯+Uk≤1. Naturally En⊂En−1 . ThenP(N=n,n≥2)=P(En−1∩Enc)=P(En−1)−P(En−1∩En)=1(n−1)!−1n!=n−1n!.
ThusE[N]=∑n=2∞nP(N=n)=∑n=2∞1(n−2)!=e.


Can also see this note from Statistics Theory on Bayes estimation.

1 Bayesian Inference

X: observed data. Θ: unknown parameter(s). All continuous random variables.
Law of Total Probability:fX(x)=∫−∞+∞fx|Θ=θ(x)fΘ(θ)dθ.
Bayes Rule for Continuous RV:fΘ|X=x(θ)=fX|Θ=θ(x)⋅fΘ(θ)∫−∞+∞fX|Θ=θ(x)fΘ(θ)dθ.
If X or Θ is discrete, use p.m.f instead of p.d.f.

Example (Binomial)

(X|Θ=θ)∼Binomial(n,θ). Θ∼Beta(α,β), (Θ|X=x)∼Beta(α+x,β+n−x). ThenP(X=x|Θ=θ)=(nx)θx(1−θ)n−x1{x∈{0,1,⋯,n}}.
ThenfΘ|X=x(θ)∝P(X=x|Θ=θ)fΘ(θ).If fΘ(θ)∝θα−1(1−θ)β−1, thenfΘ|X=x(θ)∝θα+x−1(1−θ)β+n−x−1.

For d dimensions, X→=(X1,⋯,Xd),θ→=(θ1,⋯,θd), ∑i=1dθi=1. (X→|Θ→=θ→)∼Multinomial(n,θ1,⋯,θd).P(X→=x→|Θ→=θ→)=(nx1,⋯,xd)θ1x1⋯θdxd1{∑i=1nxi=n}∏i=1d{xi∈{0,⋯,n}}.

Example (Gaussian)

X1,⋯,Xn∼i.i.dN(μ,σ2).

  1. μ unknown, σ2 known.
    X→=(X1,⋯,Xn). Random variable M=μ.fX→|M=μ(x→)∝exp⁡{−12σ2∑i=1n(xi−μ)2}.
    Conjugate prior M∼N(μ0,σ02).fM(μ)∝exp⁡{−12σ02(μ−μ02)}.
fM|X→=x(μ)∝fX→|M=μ(x→)fM(μ)∝exp⁡{−12σ2∑i=1n(xi−μ)2−12σ02(μ−μ0)2}∝exp⁡{−12σ02(μ−μn)2},

whereμn=(σ2nσ02+σ2)μ0+(nσ02nσ02+σ2)μML⏞1n∑i=1nXi,1σn2=1σ02+nσ2.
(Precision = prior precision + data precision)

  1. μn→μML as n→∞.
  2. Precisions are additive.
  3. Precision gets large as sample size gets large.
  4. For a finite n, if σ02→∞, then μn→μML and σn2→σ2n.

  1. μ known, σ2 unknown.
    Put a prior on precision Λ=1σ2. ThenfX→|Λ=λ(x→)=(λ2π)n2exp⁡{−λ2∑i=1n(xi−μ)2}.Conjugate prior Λ∼Gamma(α0,β0):fΛ(λ)=β0α0Γ(α0)λα0−1e−β0λ.fΛ|X→=x→(λ)∝fX→|Λ=λ(x→)fΛ(λ)∝λα0+n2−1exp⁡{−λ[β0+12∑i=1n(xi−μ)2]}, σML2=1n∑i=1n(xi−μ)2. Then (Λ|X→=x→)∼Gamma(αn,βn). αn=α0+n2,βn=β0+n2σML2.

  1. Both μ,σ2 unknown.fX→|M=μ,Λ=λ(x→)∝[λ12e−λμ22]nexp⁡{λμ∑i=1nxi−λ2∑i=1nxi2}. fM,Λ(μ,λ)=fM|Λ=λ(μ)⏟N(μ0,1cλ)fΛ(λ)⏟Gamma(α,β), whereμ0=ac,α=1+c2,β=b−a22c,a,b,c>0.

2 Model Selection

Double Exponential/Laplace

X∼Laplace(μ,β),β>0.fX(x)=12βexp⁡(−|x−μ|β). E[X]=μ,Var(X)=2β2.

Θ∼Bernoulli(12). Consider Model 0 vs. Model 1.
Prior odds: P(Θ=0)P(Θ=1).
X1,⋯,Xn|Θ=0∼i.i.dN(0,1), p.d.f. f0(x)=12πe−x22.
X1,⋯,Xn|Θ=1∼i.i.dLaplace(0,π2), p.d.f. f1(x)=12πe−|x|2π.
Bayes Factor: fX→|Θ=0(x→)fX→|Θ=1(x→).
Posterior odds: BF×Prior odds.

Now suppose

X1,⋯,Xn|Θ=0∼i.i.dN(0,α2),α>0 unknown.X1,⋯,Xn|Θ=1∼i.i.dLaplace(0,β),β>0 unknown.

Put priors on α,β and get RV: A,B.

For example, log⁡A|Θ=0∼Uniform(−c,c),log⁡B|Θ=1∼Uniform(−c,c).

fX→|Θ=0(x→)=∫0+∞fX→|Θ=0,A=α(x→)fA|Θ=0(α)dα.

Z=log⁡A,A=T(Z)=eZ. fA|Θ=0(α)=flog⁡A|Θ=0(log⁡α)|ddαlog⁡α|=1αflog⁡A|Θ=0(log⁡α)=1{−c<log⁡α<c}2cα,fX→|Θ=1(x→)=∫0∞fX→|Θ=1,B=β(x→)fB|Θ=1(β)dβ.fB|Θ=1(β)=1{−c<log⁡β<c}2cβ,BF=fX→|Θ=0(x→)fX→|Θ=1(x→).
(c cancels out in the Bayes Factor.)