0 Summary on Common Distributions

Name Distribution E Var MGF Char[1]
Bernoulli(p) P(X=1)=p,P(X=0)=1−p p p(1−p) 1−p+pet 1−p+peit
Binomial(n,p) P(Sn=k)=(nk)pk(1−p)n−k np np(1−p) (1−p+pet)n (1−p+peit)n
Multinomial
Geometric(p) P(W=k)=(1−p)k−1p,k∈N∗ 1p 1−pp2 pet1−(1−p)et peit1−(1−p)eit
NB(r,p)[2] P(Fr=k)=(r+k−1k)pr(1−p)k r(1−p)p r(1−p)p2 (pet1−(1−p)et)r (p1−eit+peit)r
Hypergeom(N,B,n)[3] P(X=k)=(Bk)(N−Bn−k)(Nn) nBN nBN(1−BN)(N−nn−1)

Poisson(λ)
P(X=k)=λkk!e−λ λ λ eλ(et−1) eλ(eit−1)
Uniform(a,b) f(x)=1b−a1[a,b] b+a2 (b−a)212 etb−etat(b−a) eitb−eitait(b−a)

Laplace(μ,b)
f(x)=12bexp⁡(−∣x−μ∣b) μ 2b2 etμ1−b2t2 eitμ1+b2t2
N(μ,σ2) f(x)=12πσe−(x−μ)22σ2 μ σ2 etμ+12σ2t2 eitμ−12σ2t2
χk2 (1−2t)−k2 (1−2it)−k2
Gamma(α,β) f(x)=βαΓ(α)xα−1e−βx,x>0 αβ αβ2 (1−tβ)−α (1−itβ)−α
Exp(λ) λe−λx,x≥0 1λ 1λ2 (1−tλ−1)−1 (1−itλ−1)−1
Beta(α,β) f(x)=Γ(α,β)Γ(α)Γ(β)xα−1(1−x)β−1 αα+β αβ(α+β)2(α+β+1) 1+∑k=1∞(∏r=0k−1α+rα+β+r)tkk! 1F1(α;α+β;it)
N(μ,Σ) μ Σ etT(μ+12Σt) etT(iμ−12Σt)

1 Preliminary

1.1 Generalized Binomial Coefficient

For a∈C and any integer k≥0, the generalized binomial coefficient is defined as (ak)=a(a−1)(a−2)⋯(a−k+1)k!.
Specially, if a=−m where m∈N∗, then (−mk)=(−1)k⋅(m+k−1k).

1.2 Newton Binomial Theorem

Let a∈C and x∈C with |x|<1, then

(1+x)a=∑k=0∞(ak)xk.

2 Bernoulli Trial and Bernoulli distribution

A Bernoulli trial is a random experiment having 2 possible outcomes commonly labeled as success and failure. If X is the indicator of success, then X follows a Bernoulli distribution or X∼Bernoulli(p) if P[X=1]=p,P[X=0]=1−p.

2.1 Expectation and variation of Bernoulli distribution

Assume that X∼Bernoulli(p), therefore E[X]=pE[X2]=pVar[X]=E[X2]−E[X]2=p(1−p).

3 Binomial Distribution

3.1 Introduction of binomial distribution

Assume that you are testing the toys manufactured by a factory, where the probability that a toy is defective is p. In order to decide whether to accept the toys, you randomly sample n of them from the whole batch. Let X denotes the number of broken toys (0≤X≤n). Then X follows a binomial distribution with parameters n and p X∼Bin(n,p). The pmf of X is P[X=k]=(nk)pk(1−p)n−kk∈{0,1,⋯,n}.

Tip

We can consider Binomial distribution as n independent Bernoulli trials and the random variable X∼Bin(n,p) denotes the number of successes in n independent Bernoulli trials: X=ξ1+ξ2+⋯+ξn,ξi∼Bernoulli(p).

3.2 Expectation and Variation of binomial distribution

  1. We can directly compute the expectation and variation by calculating the first and second moment of X. E[X]=∑k=0nk⋅(nK)pk(1−p)n−k=∑k=1nk⋅(nK)pk(1−p)n−k=np∑k=1n(n−1K−1)pk−1(1−p)(n−1)−(k−1)=np∑t=0n−1(n−1T)pt(1−p)n−1−t=np.
E[X2]=∑k=0nk2(nK)pk(1−p)n−k=∑k=1nk(k−1)(nK)pk(1−p)n−k+∑k=1nk(nK)pk(1−p)n−k=∑k=2nk(k−1)(nK)pk(1−p)n−k+E[X]=p2n(n−1)∑k=2n(n−2K−2)pk−2(1−p)(n−2)−(k−2)+E[X]=p2n(n−1)∑t=0n−2(n−2T)pk−2(1−p)k−2−t+E[X]=p2n(n−1)+np.

therefore E[X]=npVar[X]=E[X2]−E[X]2=np(1−p).
2. Let ξi denote the indicator variable of success of the ith Bernoulli trial. Then the number of successes can be expressed as X=ξ1+⋯+ξn. Using the expectation and variation of Bernoulli distribution and linearity of expectation and variation for independent random variables. We get E[X]=E[ξ1+⋯+ξn]=∑i=1nE[ξi]=npVar[X]=Var[ξ1+⋯+ξn]=∑i=1nVar[ξi]=np(1−p).

4 Geometric Distribution

4.1 Pmf of Geometric Distribution

The geometric distribution gives the probability that the first occurrence of success requires k independent trials, each with probability p. Assume a random variable X∼Geo(p), then P[X=k]=(1−p)k−1p,k=1,2,3,⋯.

4.2 Expectation of Geometric Distribution

Here we use an interesting trick to solve E[X] where X∼Geo(p). E[X]=∑k=1∞k⋅(1−p)k−1p=p∑k=1∞[(q)k]′=p⋅(q1−q)′=1p.

4.3 Variation of Geometric Distribution

4.4 Sum of Geometric-distributed random variables

Assume we have n random variables X1,⋯,Xn satisfying Xi∼Geo (p), i=1,⋯,n. From the definition of Geometric distribution we know that it is about the number of trials when the first success occurs. Then Y=X1+X2+⋯+Xn indicates the number of trials when the nth success occurs. Therefore the last trial has succeeded and the remaining trials have n−1 successes. Thus P[Y=k]=p⋅(k−1n−1)pn−1(1−p)k−n=(k−1n−1)pn(1−p)k−n.We call this distribution Negative Binomial Distribution.

5 Pascal Distribution

In n independent Bernoulli trials with success probability p, the pascal distribution focuses on the number of failures when the rth success happens. If a random variable X∼Pascal(r,p), then P[X=k]=(k+r−1k)pr(1−p)k,k=0,1,2,⋯.

5.1 Expectation and variation of Pascal distribution

If X∼Pascal(r,p), then E[X]=r(1−p)p,Var[X]=r(1−p)p2.

6 Negative Binomial Distribution

In n independent Bernoulli trials with success probability p, the negative binomial distribution focuses on the number of trials when the rth success happens. If a random variable X∼NB(r,p), then P[X=k]=(k−1r−1)pr(1−p)k−r,k=r,r+1,⋯.From the above discussion about geometric distribution, we have X=X1+X2+⋯Xr where Xi∼Geo(p)(i∈[r]), where Xi denotes the number of trials needed after the (i−1)st success to obtain the ith success. (X1,⋯,Xr are independent)

6.1 Expectation and variation of negative binomial distribution

E[X]=∑k=r∞k⋅(k−1R−1)pr(1−p)k−r

7 Gaussian Distribution

7.1 Pdf of Gaussian Distribution

If a random variable X∼N(μ,σ2), where μ is the mean and σ2 is the variance, then the pdf of X is f(x)=12πσexp⁡{−(x−μ)22σ2}.

7.2 Mgf of Gaussian Distribution

Assume a random variable X∼N(μ,σ2), then the mgf of X is MX(t)=E[etX]=exp⁡(μt+12σ2t2).

7.3 Characteristic function of Gaussian Variables

Since we know the mgf of Gaussian variable X∼N(μ,σ) is MX(t)=E[etX]=exp⁡(μt+12σ2t2),therefore the characteristic function of X is equal to φX(t)=E[eitX]=MX(it)=exp⁡(iμt−12σ2t2).

7.4 Moments of standard Gaussian

Assume that the random variable X∼N(0,1) and that k∈N∗, then E[Xk]={0k odd(k−1)!!k even.
First we state an important lemma:

Relation between Moments and derivatives of mgf

MX(n)(0)=E(Xn)

Since we have the mgf of Gaussian variables, moments can be obtained via differentiation. We also provide a direct proof for the moments of the standard Gaussian.

If k is odd, then it's obvious that E[Xk]=0, so the following process assumes that k is even. Then 12π∫−∞∞xke−x22dx=2π∫0∞xke−x22dx=2π∫0∞tk−12e−t2dt=2π⋅12⋅Γ(k+12)(12)(k+1)/2=12π2k+12⋅(k−1)!!2k2⋅Γ(12)=(k−1)!!(Recall that Γ(12)=π).

7.5 An important property of Standard Gaussian and Mills Ratio

A standard Gaussian variable X∼N(0,1) has a good property in terms of derivatives that ϕ′(x)=−xϕ(x).

7.6 Expectation of absolute value of Gaussian variables

If a random variable Z∼N(0,1), E[|Z|]=2π.

7.7 Rotational Invariance of Gaussian Variables

If a matrix R∈Rn×n is orthogonal, indicating that RR⊤=R⊤R=In, given a random vector X∼N(0,σ2In), it satisfies RX=dY.This can be proved using linear transformation of multivariate Gaussian variables.

8 Multivariate Gaussian Distribution

8.1 Definition of Standard Normal Random Vector

A real random vector X=(X1,⋯,Xk)⊤ is called a standard normal random vector if all of its components Xi (i∈[k]) are independent standard Gaussian variables. We denote it X∼N(0,In) where the mean vector is 0→ and the covariance matrix is In.

8.2 Definition of Normal random vector

A real random vector X=(X1,⋯,Xk)⊤ is called a normal random vector if there exists a random normal random vector Z∈Rl, a k×l matrix A and a k-dim vector μ, such that X=AZ+μ. (Wikipedia) We denote it X∼N(μ,Σ), where μ is the mean vector and Σ is the covariance matrix with μ=(μ1μ2⋮μk),Σ=(σ12ρ12σ1σ2⋯ρ1kσ1σkρ21σ2σ1σ22⋯ρ2kσ2σk⋮⋮⋯⋮ρk1σkσ1ρk2σkσ2⋯σk2),where Σij is the covariance between Xi and Xj (i,j∈[k]) and μi is the mean of Xi (i∈[k]).

8.3 Joint pdf of multivariate Gaussian distribution

Assume that X∈Rk∼N(μ,Σ), then fX(x1,⋯,xk)=1(2π)k|Σ|exp⁡{−12(X−μ)⊤Σ−1(X−μ)}.

8.4 Characteristic function of multivariate Gaussian variables

If X∈Rk∼N(μ,Σ), then φX(t)=E[eit⊤X]=exp⁡{it⊤μ−12t⊤Σt},t∈Rk.

8.5 Linear Transformation of multivariate Gaussian variables

If X∼N(μ,Σ) and Y=α+AX, then E[Y]=Aμ+αVar[Y]=AΣA⊤.

9 Chi-squared Distribution

9.1 Definition of Chi-squared Distribution

Assume that we have n i.i.d. samples X1,⋯,Xn from N(0,1). Then ∑i=1nXi2∼χn2,which is called Chi-squared Distribution with n degrees of freedom.

9.2 Pdf of Chi-squared Distribution

From the first part, if X∼N(0,1), then Y=X2∼χ12. Then FY(y)=P[X2≤y]=P[−y≤X≤y]=∫−yy12πe−t22dt.Let G(t) denote the primitive function of 12πexp⁡{−t22}. Therefore fY(y)=FY′(y)=(G(y)−G(−y))′=12πe−y21y(1)=(12)1/2Γ(12)y12−1e−12ywhere (1) is because Γ(12)=π.

pdf of χ12

fY(y)=12πe−y21y,y≥0

9.3 Relation to Gamma Distribution

In the derivation of the pdf of Chi-squared Distribution we can conclude that χ12⇔Γ(12,12),which is a special case of Gamma distribution. Therefore, if a random variable X∼χn2, then X∼Γ(n2,12).

9.4 Expectation of Chi-squared Distribution

Given that if a random variable Y∼Γ(α,β), then E[Y]=αβ, therefore for a random variable X∼χn2, E[X]=n.

9.5 Variation of Chi-squared Distribution

Given that if a random variable Y∼Γ(α,β), then Var[Y]=αβ2, therefore for a random variable X∼χn2, Var[X]=2n.

9.6 Mgf of Chi-squared Distribution

Given that for a Gamma distributed random variable X∼Γ(α,β), the mgf is MX(t)=1(1−tβ)α,since Y∼χn2 is equivalent to Y∼Γ(n2,12), therefore MY(t)=1(1−2t)n2.

9.7 Characteristic function of Chi-squared Distribution

Since for Y∼χn2, MY(t)=(1−2t)−n2, therefore φY(t)=MY(it)=(1−2it)−n2.

9.8 Asymptotic Property of Chi-squared Distribution

9.8.1 LLN

From the definition of Chi-squared Distribution we have Y=∑i=1nXi2∼χn2,Xi∼N(0,1).If we treat every Xi2 as an i.i.d. sample from the same distribution with mean E[Xi2]=1. By the law of large number we can derive ∑i=1nXi2n→PE[Xi2]=1,thus Yn→P1,Y∼χn2.

9.8.2 CLT

Similar to the LLN part, since we treated every Xi2 as an i.i.d. sample, therefore we can apply Central Limit Theorem to get convergence in distribution:Y−n2n→LN(0,1),Y∼χn2.

10 Exponential Distribution

10.1 Pdf of Exponential Distribution

If a random variable X∼Exp(λ), then the pdf of X is f(x)={λe−λxx>00elsewhere.

10.2 The tail probability of Exponential Distribution

If X∼Exp(λ) with λ>0, then the tail probability is P[X≥t]=e−λt.

10.3 Memoryless Property

Memoryless property of Exponential distributed variable

P[X>s+t∣X>s]=P[X>t].

Since P[X>s+t∣X>s]=P[X>s+t]P[X>s]=e−λ(s+t)/e−λs=e−λt.
Changing the expression P[X>s+t∣X>s] to P[X−s>t∣X>s] we can conclude that X−s∣X>s and X have the same distribution.

10.4 Relation to Gamma Distribution

Exponential Distribution is a special form of Gamma Distribution, using shape-rate version we can find that if X∼Γ(1,α):f(x)=αΓ(1)x1−1e−αx=αe−αx,x>0.Therefore X∼Exp(λ)⇔X∼Γ(1,λ).

11 Poisson Distribution

11.1 Pmf of Poisson Distribution

Assume random variable X∼Poisson(λ), the pmf of X is given as follows:P[X=k]=e−λλkk!λ>0,k=0,1,2,⋯.

11.2 Expectation and variance of Poisson Distribution

First we have an important property of Poisson Distribution:Assume X∼Poisson(λ), then E[X(X−1)⋯(X−k)]=λk+1,k=0,1,2,⋯Using this property we can quickly get that E[X]=λE[X2]=λ2+λ,therefore Var[X]=E[X2]−E[X]2=λ.

11.3 Mgf of Poisson Distribution

If a random variable X∼Poisson(λ) where λ>0, then the mgf of X is MX(t)=exp⁡{λ(et−1)}.

11.4 Reproducibility of Poisson distribution

If X1,⋯,Xn are independent variables satisfying Xi∼Poisson(λi), then ∑i=1nXi∼Poisson(∑i=1nλi).

11.5 Poisson approximation to the Binomial distribution

For n Bernoulli trials with probability p, if n is large enough and p is small enough, then the Binomial distribution is approximately equal to Poisson distribution with parameter λ=np.

12 Gamma Distribution

12.1 Γ function

The Gamma function Γ(⋅) is defined as follows:

Γ(α)=∫0∞yα−1e−ydy,

An integration by parts shows thatΓ(α)=(α−1)∫0∞yα−2e−ydy=(α−1)Γ(α−1)Given that Γ(1)=∫0∞e−ydy=1Thus if α is a positive integer greater than 1, Γ(α)=(α−1)!

Recursive property of Gamma function

For all x>0, the Gamma function satisfies the following recursion:Γ(x+1)=xΓ(x),x>0.

12.2 Γ(α,β) Distribution (shape-scale version)

We say that the continuous random variable X has a Γ−distribution with parameters α>0 and β>0 if its pdf is f(x)={1Γ(α)βαxα−1e−x/β0<x<∞0elsewhere,we often write that Γ(α,β) distribution where α is the shape parameter and β is the scale parameter.

12.3 Γ(α,β) Distribution (shape-rate version)

We say that the continuous random variable X has a Γ−distribution with parameters α>0 and β>0 if its pdf is f(x)={βαΓ(α)xα−1e−βx0<x<∞0elsewhere,we often write that Γ(α,β) distribution where α is the shape parameter and β is the rate parameter. Throughout this article we use this version of Gamma distribution.

12.4 Expectation of Γ(α,β) Distribution

We can use the definition of the Gamma function to simplify the computation of the integral:E[X]=∫0∞βαΓ(α)xα−1e−βx⋅xdx=∫0∞βαΓ(α)xαe−βxdx=βαΓ(α)⋅Γ(α+1)βα+1∫0∞βα+1Γ(α+1)xαe−βxdx=αβ.

12.5 Variation of Γ(α,β) Distribution

Similarly, we can use the definition of the Gamma function to simplify the computation of the integral:E[X2]=∫0∞βαΓ(α)xα−1e−βx⋅x2dx=∫0∞βαΓ(α)xα+1e−βxdx=βαΓ(α)⋅Γ(α+2)βα+2∫0∞βα+2Γ(α+2)xα+1e−βxdx=α(α+1)β2.Therefore Var[X2]=E[X2]−(E[X])2=αβ2.

12.6 Mgf of Γ(α,β) distribution

MX(t)=∫0∞etx⋅βαΓ(α)xα−1e−βxdx=∫0∞βαΓ(α)xα−1e−(β−t)xdx=(11−tβ)α.

12.6.1 Additivity property of the Gamma Distribution

If X∼Γ(α1,θ) and Y∼Γ(α2,θ) are independent, then X+Y∼Γ(α1+α2,θ).

This can be proved using mgf of Gamma distribution.

13 Inverse Gamma Distribution

13.1 Pdf of Inverse Gamma Distribution

If a random variable Y∼IG(α,β), then f(y)=βαΓ(α)y−(α+1)e−β/y,y>0.

14 Beta Distribution

14.1 Beta Function

Beta function is defined by integralB(r1,r2)=∫01tr1−1(1−t)r2−1dt.

Association between Beta function and Gamma function

B(α,β)=Γ(α)Γ(β)Γ(α+β).

Apply the variable substitution by letting x=st, y=s(1−t), then the above integral is equivalent to ∫0∞∫01e−s(st)α−1(s(1−t))β−1sdtds=∫0∞e−ssα+β−1ds⋅∫01tα−1(1−t)β−1dt=Γ(α+β)⋅B(α,β),thus we have finished the proof.

Another expression of Beta function

We can get another version of Beta function using variable substitution t=x1+x,x∈(0,∞). Then the integral is equivalent to B(r1,r2)=∫0∞(x1+x)r1−1(11+x)r2−11(1+x)2dx=∫0∞xr1−1(11+x)r1+r2dx=∫0∞xr1−1(1+x)−r1−r2dx.

14.2 Pdf of beta distribution

The beta distribution beta(α,β) is a two-parameter distribution with range [0,1] and pdf f(x)=(α+β−1)!(α−1)!(β−1)!xα−1(1−x)β−1(α>0,β>0)=1B(α,β)xα−1(1−x)β−1.

14.3 Expectation of beta distribution

We can use the pdf of beta distribution to get the expectation of a beta distribution random variable easily, as the computation that follows:E[X]=∫011B(α,β)xα−1(1−x)β−1⋅xdx=∫011B(α,β)xα(1−x)β−1dx=B(α+1,β)B(α,β)∫011B(α+1,β)xα(1−x)β−1dx=B(α+1,β)B(α,β)=αα+β.

14.4 Variation of beta distribution

Similarly, we can use the pdf of beta distribution to derive E[X2], using Var[X]=E[X2]−E[X]2 we can get the variation.E[X2]=B(α+2,β)B(α,β), Var[X]=E[X2]−E[X]2=B(α+2,β)B(α,β)−(B(α+1,β)B(α,β))2=αβ(α+β)2(α+β+1).

14.5 Relation to Chi-squared Distribution

Let X and Y are independent variables and satisfy X∼χn2,Y∼χm2. Then we have XX+Y∼B(n2,m2).

15 Cauchy Distribution

15.1 Pdf of Cauchy Distribution

A random variable X is said to follow a Cauchy Distribution with location parameter θ∈R and scale parameter γ>0 if f(x;θ,γ)=1πγ[1+(x−θγ)2],x∈R.Therefore X∼Cauchy(θ,γ).

A special property of Cauchy Distribution

The expectation of Cauchy Distribution doesn't exist.

15.2 Median of Cauchy Distribution

15.3 Characteristic function of Cauchy distribution

16 T Distribution

16.1 Definition

Assume that X∼N(0,1) and Y∼χn2 are independent. Then the statistic

T=XY/n

Is said to follow a t-distribution with n degrees of freedom.

16.2 Pdf of t-distribution

For a random variable T∼tn, the density function of T is given by fT(t)=Γ(n+12)Γ(n2)πn(t2n+1)−n+12.

17 F Distribution

17.1 Definition

Assume that X∼χn2 and Y∼χm2 are independent. Then the statisticF=X/nY/m∼Fn,mis said to follow a F-distribution with parameters n and m.

17.2 Pdf of F-distribution

From the definition we have F=X/nY/m, this is an example of the probability density function of the ratio of two Chi-squared distributed variables. The pdf of F∼Fn,m is given bellow:fF(f)=Γ(m+n2)Γ(m2)Γ(n2)(nm)n2fn2−1(nmf+1)−m+n2.

17.3 Property of F-distribution

Fn,m(1−α)=1Fm,n(α).

Thus we have finished the proof.

17.4 Expectation and Variation of F-distribution

If a random variable X∼Fn,m, then E[X]=mm−2,Var[X]=2m2(n+m−2)n(m−2)2(m−4).

For variation of F-distribution we can use similar calculation as its expectation using the second version of Beta function.


  1. Characteristic function ↩︎

  2. NegativeBinomial ↩︎

  3. HyperGeometric ↩︎