11 AR Models

1 AR Models

AR Model

The AR (Auto-Regression) model of order p (AR(p)) is given by (1.1)yt=ϕ0+ϕ1yt−1+⋯+ϕpyt−p+εt for t=p+1,⋯,n. In matrix notation, Y=Xβ+ε, where Y=(yp+1⋮yn),X=(1yp⋯y1 ⋮⋮⋱⋮1yn−1⋯yn−p),β=(ϕ0⋮ϕp),ε=(εp+1⋮εn).

yt is regressed on its own lagged values yt−1,⋯,yt−p.

For predicting yn+1, plug t=n+1 into (1.1): yn+1=ϕ^0+ϕ^1yn+⋯+ϕ^pyn+1−p.
Note that yn,yn−1,⋯,yn+1−p are all observed and they are the last p observations. Then for yn+2, yn+2=ϕ^0+ϕ^1yn+1+⋯+ϕ^pyn+2−p.
Here yn+1 is not observed, but we can replace with predicted value y^n+1. And recursively, predict value for any time point.

We can see AR as simply regression of the observed time series on lagged versions of itself.

2 Relation to MLE

Start with AR(1): (2.1)yt=ϕ0+ϕ1yt−1+εt,t=2,⋯,n.

2.1 Usual Regression

(2) looks just like (2.2)yi=β0+β1xi+εi. For this, given independence, Likelihood=fx1,y1,⋯,xn,yn|θ(x1,y1,⋯,xn,yn)=∏i=1nfxi,yi|θ(xi,yi)=∏i=1nfyi|xi,θ(yi)fxi|θ(xi).
Now replace yi by β0+β1xi+εi: Likelihood=∏i=1nfβ0+β1xi+εi|xi,θ(yi)fxi|θ(xi)=∏i=1nfεi|xi,θ(yi−β0−β1xi)fxi|θ(xi)=∏i=1mfεi|θ(yi−β0−β1xi)fxi|θ(xi)=∏i=1m12πσexp⁡(−(yi−β0−β1xi)2σ2)fxi|θ(xi)=(12πσ)mexp⁡(−12σ2∑i=1m(yi−β0−β1xi)2)∏i=1mfxi|θ(xi)∝(12πσ)mexp⁡(−12σ2∑i=1m(yi−β0−β1xi)2).
To sum up, here are the assumptions we made:

2.2 Back to AR(1)

When we try to apply (3) to (2), these assumptions don't hold:

So now given y1,⋯,yn, Likelihood for (2)=fy1,⋯,yn|θ(y1,⋯,yn)=fy1|θ(y1)fy2|y1,θ(y2)⋯fyn|y1,⋯,yn−1,θ(yn)=fy1|θ(y1)∏t=2nfyt|y1,⋯,yt−1,θ(yt)=fy1|θ(y1)∏t=2nfεt|y1,⋯,yt−1,θ(yt−ϕ0−ϕ1yt−1)=fy1|θ(y1)∏t=2nfεt(yt−ϕ0−ϕ1yt−1)=fy1|θ(y1)∏t=2n12πσexp⁡(−12σ2(yt−ϕ0−ϕ1yt−1)2)=fy1|θ(y1)(12πσ)n−1exp⁡(−12σ2∑t=2n(yt−ϕ0−ϕ1yt−1)2).
Here we assume

2.2.1 Computation of fy1|θ(y1)

We can't get fy1|θ(y1) from equation. There are two approaches to get it.

First, we simply assume fy1|θ(y1) does not depend on θ. So (2.3)Likelihood∝(12πσ)n−1exp⁡(−12σ2∑t=2n(yt−ϕ0−ϕ1yt−1)2).
It's easy to verify that the estimates ϕ^0,ϕ^1 are identical to those obtained in (2.1). We call this Conditional MLE.

Second, we extend the model to t=1,0,−1,−2,⋯. So y1=ϕ0+ϕ1y0+ε1=ϕ0+ϕ1(ϕ0+ϕ1y−1+ε0)+ε1=ϕ0(1+ϕ1)+ϕ12y−1+ϕ1ε0+ε1=ϕ0(1+ϕ1)+ϕ12(ϕ0+ϕ1y−2+ε−1)+ϕ1ε0+ε1=ϕ0(1+ϕ1+ϕ12)+ϕ13y−2+ϕ12ε−1+ϕ1ε0+ε1=⋯=ϕ0∑j=0Mϕ1j+ϕ1M+1y−M+∑j=0Mϕ1jε1−j.
If |ϕ1|<1, coefficient ϕ1M+1 is very small, so y1≈ϕ0∑j=0Mϕ1j+∑j=0Mϕ1jε1−j≈ϕ0∑j=0∞ϕ1j+∑j=0∞ϕ1jε1−j=ϕ01−ϕ1+∑j=0∞ϕ1jε1−j.
Since E(∑j=0∞ϕ1jε1−j)=0, and Var(∑j=0∞ϕ1jε1−j)=∑j=0∞Var(ϕ1jε1−j)=∑j=0∞ϕ12jVar(ε1−j)=σ2∑j=0∞ϕ12j=σ21−ϕ12.
Thus when |ϕ1|<1, y1∼N(ϕ01−ϕ1,σ21−ϕ12), so fy1|θ(y1)=1−ϕ122πσexp⁡(−1−ϕ122σ2(y1−ϕ01−ϕ1)2), so finally (5)Likelihood=1−ϕ12(2πσ)nexp⁡(−1−ϕ122σ2(y1−ϕ01−ϕ1)2−12σ2∑t=2n(yt−ϕ0−ϕ1yt−1)2).
Now compare (4) and (5). (5) is referred to as full likelihood for AR(1) and therefore full MLE. (4) and (5) will be quite close when |ϕ1|<1 and n is large.

2.3 AR(p)

The AR(p) model is given by yt=ϕ0+ϕ1yt−1+⋯+ϕpyt−p+εt.
The likelihood is fy1,⋯,yn|θ(y1,⋯,yn)=fyp+1,⋯,yn|y1,⋯,yp,θ(yp+1,⋯,yn)fy1,⋯,yp|θ(y1,⋯,yp).
The conditional likelihood is fyp+1,⋯,yn|y1,⋯,yp,θ(yp+1,⋯,yn)=∏t=p+1nfyt|yt−1,⋯,y1(yt)=∏t=p+1nfεt|yt−1,⋯,y1(yt−ϕ0−ϕ1yt−1−⋯−ϕpyt−p)=(12πσ)n−pexp⁡(−12σ2∑t=p+1n(yt−ϕ0−ϕ1yt−1−⋯−ϕpyt−p)2).
Here we assume εt|yt−1,⋯,y1∼N(0,σ2),t=p+1,⋯,n.

To obtain the parameter estimates, we can directly maximize the likelihood. And since fy1,⋯,yp|θ(y1,⋯,yp) does not depend on θ, it is equivalent to maximizing the conditional likelihood.

2.3.1 Bayesian Approach

However, if we want to derive fy1,⋯,yp|θ(y1,⋯,yp) in a more principled way, we have to use (1.1) for smaller values of t but it's complicated and not really worth it. We can also work under some "stationarity" assumptions on ϕ0,⋯,ϕp (much simpler than conditional likelihood): use matrix notation (see here), and likelihood∝θ(12πσ)n−1exp⁡(−||Y−Xβ||22σ2). Assume ϕ0,⋯,ϕp,log⁡σ∼i.i.dUnif(−C,C), then by here, β|data∼tn−2p−1,p+1(β^,σ^2(XTX)−1), where β^=(XTX)−1XTY,σ^=||Y−Xβ^||2n−2p−1. If inference for σ is desired, we can use ||Y−Xβ^||2σ2|data∼χn−2p−12.

Bayesian inference for AR is identical to linear regression models because of the same likelihood. Bayesian inference only cares about the likelihood.
Frequentist inference is based on MLE, given by β^ and σ^MLE=||Y−Xβ^||2n−p. The analysis is quite different from linear regression. The results are slightly different but close.

3 Predictions and Difference Equations

Given a fitted AR(p) model with ϕ^0,⋯,ϕ^p, predictions y^n+i for i=1,2,⋯ are obtained by: (3.1)y^n+i=ϕ^0+ϕ^1y^n+i−1+⋯+ϕ^py^n+i−p,i=1,2,⋯ where the recursion is initialized with y^j=yj,j=n,n−1,⋯,n+1−p.
Or we can rewrite to (3.2)uk=α0+α1uk−1+⋯+αpuk−p,k=p,p+1,⋯ (Initialized by u0,⋯,up−1). (3.2) is called a difference equation of order p.

3.1 First Order (p=1)

Now (3.2) becomes uk=α0+α1uk−1 along with initialized u0. Convert to a homogeneous equation (with no intercept term) by taking vk=uk−α01−α1: vk=α1vk−1.
So vk=α1kv0⇒uk=1−α1k1−α1α0+α1ku0.

3.2 General Case: Bayesian Approach

In the Bayesian context, prediction is done via joint probability distribution of yn+1,⋯,yn+k conditional on y1,⋯,yn. Now consider conditional expectations: (3.3)E(yn+i|y1,⋯,yn)=∫E(yn+i|y1,⋯,yn,θ)fθ|y1,⋯,yn(θ)dθ.
First calculate y^n+i(θ)=E(yn+i|y1,⋯,yn,θ) for fixed θ: (3.4)y^n+i(θ)=ϕ0+ϕ1y^n+i−1(θ)+⋯+ϕpy^n+i−p(θ).
If we initialize this with y^j(θ)=yj, j=n,n−1,⋯,n+1−p, then (3.4) can be evaluated in sequence for i=1,2,⋯.

Now (3.3) becomes E(yn+i|y1,⋯,yn)=∫y^n+i(θ)fθ|y1,⋯,yn(θ)dθ.
We can do one of two things to compute:

  1. Generate posterior samples θ(1),⋯,θ(N) from fθ|y1,⋯,yn(θ), then E(yn+i|y1,⋯,yn)≈1N∑j=1Ny^n+i(θ(j)).
  2. Use the fact that fθ|y1,⋯,yn(θ) is usually highly concentrated around θ^=(β^,σ^). We then ignore the small uncertainty of θ around θ^: E(yn+i|y1,⋯,yn)≈y^n+i(θ^).