1.1 Probability

We have defined probability in STAT 201A note. However for a different course with (possibly) different notations, we will derive again.

1 What is Probability?

Probability

Mathematical probability is a function P mapping some subsets of a sample space X to [0,1], satisfying

  • P(X)=1,
  • P(∅)=0,
  • (disjoint additivity) P(⋃i=1∞Ai)=∑i=1∞P(Ai),Ai∩Aj=∅,∀i≠j.

2 Measure and Integrals

2.1 Measure

Given a set X, a measure μ is a certain kind of function mapping "nice enough" subsets A⊂X to non-negative numbers μ(A)∈[0,+∞).

Generally, the domain of a measure is just a collection of "nice" subsets F⊂2X. F should satisfy certain closure properties.
First we need definition of σ− algebra. It's not important. See definition here.

Measure

Given measurable space (X,F), a measure μ is a map F→[0,+∞] with disjoint additivity and μ(∅)=0. μ is probability measure if μ(X)=1.

2.2 Integral

Integral

An integral w.r.t μ puts weight μ(A) on A. Define ∫A1{x∈A}dμ(x)=μ(A). We can extend to other functions by linearity ∫c1A(x)dμ(x)=cμ(A)⇒∫∑i=1nci1Ai(x)dμ(x)=c∑i=1nμ(Ai),
and limits ∫fdμ=limn→∞∫fn(x)dμ.
Pasted image 20241201225736.png|600

3 Densities

A measure P is absolutely continuous w.r.t μ, if P(A)=0 whenever μ(A)=0. Denote as P≪μ or μ dominates P.
If P≪μ, then define density function p:X→[0,+∞) s.t. P(A)=∫1A(x)p(x)dμ(x),∀A∈F, and by extension ∫f(x)dP(x)=∫f(x)p(x)dμ(x). Density function p is also called Radon-Nikodym derivative of P w.r.t μ, and is sometimes written as dPdμ(x).

If we don't specify μ, it is the Lebesgue measure. I.e., P is absolutely continuous in default means P≪λ.

If P is probability, and μ is Lebesgue measure, then p is called probability density function (p.d.f);
Elif μ is counting measure, p is called probability mass function (p.m.f).

4 Probability Spaces, Random Variables

Denote the outcome space Ω. To evaluate whatever P(X,⋯), it is convenient to start with abstract outcome ω∈Ω.

Probability Space

We have (Ω,F,P) a probability space.

  • ω∈Ω is called outcome.
  • A⊂Ω is called event.
  • P(A) is called probability of event A.
  • Function X:Ω→X is called random variable X(ω).
  • We say X has distribution Q (defined as X∼Q) if P(X∈B)=Q(B).

Q(B) is the push-forward of P through function X(ω): Q(B)≡P∘X−1(B).

Applied more generally, if μ is a measure on X, f:X→Y leads to new measure ν(B)⏟∈Y=μ∘f−1(B)⏟∈X. (f−1 denotes the preimage)

If P(A)=1, we say A happens almost surely.