2 Probability Basics

1 Random Variable and Distribution

Random Variable

Given a probability space (Ω,F,P). A random variable(RV) is a function X:Ω→R, s.t. X≤x≡{ω∈Ω|X(ω)≤x}∈F for all x∈R. (we call this F− measurable)

Distribution, c.d.f

Given a RV X, its (cumulative) distribution function, c.d.f FX is defined as FX(x)=P(X≤x),x∈(−∞,∞).

1.1 Discrete RV

A discrete RV means Range(X) is either finite or countably infinite.

Like, for A∈F, indicator RV IA(ω)={1,ω∈A,0,ω∉A is a discrete RV.

It has to satisfy:

1.2 Continuous RV

FX(x) is continuous, ∀x∈R. Then, we consider probability density function, p.d.f fX(x)=dFX(x)dx, and FX(x)=∫−∞xfX(y)dy.
More generally, P(X∈A)=∫AfX(x)dx.

2 Expectation

Expectation/Mean

For a function g:R→R, define expectation E.

  • For discrete RV, E[g(X)]=∑x∈Range(X)g(x)P(X=x).
  • For continuous RV, E[g(X)]=∫−∞∞g(x)fX(x)dx.

Provided that E[|g(X)|]<∞. (absolutely convergent.)

By the following theorem, we need absolute convergence to ensure E[g(X)] is well defined.

Theorem (Riemann Rearrangement Theorem)

If ∑i=1∞an converges but ∑i=1∞|an| diverges, then for any given r∈[−∞,∞], ∃ a permutation π: ∑i=1∞aπ(n)=r.

Variance

Define variance of X: Var(X)=E[(X−E[X])]2.

Claim (Linearity of Expectation)

Let X1,⋯,Xn be RVs defined on the same probability space (Ω,F,P) and E[Xi] are well defined. Then for all constants c1,⋯,cn, E[∑i=1nciXi]=∑i=1nciE[Xi].

By the claim, we can compute Var(X)=E[X2]−(E[X])2.

Covariance

Let X,Y be RVs on the same probability space. Then Cov(X,Y)=E[(X−E[X])(Y−E[Y])]=E[XY]−E[X]E[Y].

Similarly, we can show

Var(X+Y)=Var(X)+Var(Y)+2Cov(X,Y).
Theorem (Tail Sum Formula)

Let X be a RV with range {0,1,2,⋯}. Then E[X]=∑k=1∞P(X≥k).

By definition, this is easy to prove.

3 Conditional Probability

Given that event B happens, what is the probability that A also happens?
We want to consider a new probability space (B,FB,PB). How should we define PB so that it is consistent with P?

For all E1,E2∈F s.t. E1∩B≠∅,E2∩B≠∅, we want P(E1∩B)P(E2∩B)=PB(E1∩B)PB(E2∩B)⇒PB=cP.
Since PB(B)=1, we know c=P(B)−1.

So for all A,b∈F s.t. P(B)>0, define conditional probability P(A|B)=PB(A∩B)=P(A∩B)P(B).
A,B are independent if P(A|B)=P(A), i.e. P(A∩B)=P(A)P(B).

Independence of Events

Events E1,⋯,En are independent iff for every k=2,⋯,n and every k− subset {i1,⋯,ik}⊂{1,⋯,n}, P(⋂j=1kEij)=∏j=1kP(Eij).

Independence of RVs

RVs X,Y on the same probability space (Ω,F,P) are said to be independent (⊥⊥) iff P[(X≤x)∩(Y≤y)]=P(X≤x)P(Y≤y),∀x,y∈R.
Equivalent condition:

  • For discrete case, P[(X=x)∩(Y=y)]=P(X=x)P(Y=y).
  • For continuous case, fX,Y(x,y)=fX(x)fY(y).

RVs X1,⋯,Xn on the same probability space (Ω,F,P) are said to be mutually independent iff P[⋂i=1n(Xi≤xi)]=∏i=1nP(Xi≤xi),∀x1,⋯,xn∈R.

Partition

A1,⋯,An∈F. B∈F is a partition if

  • B=A1∪⋯∪An.
  • Ai∩Aj=∅,∀i≠j.
Theorem (Law of Total Probability)

Suppose B1,⋯,Bn∈F is a partition of Ω, s.t. P(Bi)>0,∀i. then ∀A∈F, P(A)=∑i=1nP(A|Bi)P(Bi).

A similar result applies to a countably infinite partition of Ω.

4 Bayes' Formula

Theorem (Bayes' Formula)

Let B1,⋯,Bn∈F be a partition of Ω. Then ∀A∈F with P(A)>0, P(Bi|A)=P(Bi∩A)P(A)=P(A|Bi)P(Bi)∑j=1nP(A|Bj)P(Bj).