# Exact Linear Rational Expectations Models

```{note}
This section is based on Lars Peter Hansen and Thomas J. Sargent, "Exact Linear Rational
Expectations Models: Specification and Estimation," Chapter 3 of *Rational Expectations
Econometrics* (Westview Press, 1991). We follow the paper's Sections 1–5, present only
Examples 1, 2, and 5, and have shortened and reorganized the exposition. The technical
identification arguments of Section 3 are summarized rather than reproduced in full.
```

This section follows {doc}`36a_interpreting_vars`, reversing the order of the two chapters in the
source. {doc}`22_rational_expectations` and {doc}`28_sims_money_income` raise the difficulty that
Chapter 4 diagnoses, so that chapter comes first here; the constructive apparatus of Chapter 3 then
answers it.

A distinguishing feature of econometric models that build in rational expectations is a set of
**cross-equation restrictions**: the parameters of the equations describing variables that
people *choose* are inherited from the equations describing the stochastic environment people
*forecast*. Because decisions depend on forecasts, and forecasts depend on the laws of motion of
the environment, the decision equations cannot have free parameters. Even for models that are
linear in the variables, these restrictions are typically nonlinear in the underlying
parameters.

This section describes a convenient way to characterize and impose those restrictions for a
class of **exact** linear rational expectations models — models in which an exact linear relation
ties forecasts of future values of one set of variables to current and past values of another
set, *and every variable entering the relation is observed by the econometrician*. That last
requirement is what makes the model "exact": the relation the econometrician confronts would hold
without error if agents could forecast perfectly. Many models of the term structure, of stock
prices, of consumption and permanent income, and of the dynamic demand for factors of production
belong to this class.

The strategy throughout is to work with a **vector moving-average representation** of the
observed process and to deduce the restrictions by straightforward applications of the
Wiener–Kolmogorov prediction formula {eq}`eq-62` — the same annihilation-operator machinery used
in {doc}`19_signal_extraction` and {doc}`18a_partial_fractions`. Once the restrictions are in
hand, the constrained moving average is nested inside a less constrained one, and the model can
be estimated and tested by maximum likelihood.

## The general model

Let $x = \{x_t\}$ be a $p$-dimensional, covariance-stationary process with mean zero. Two
information sets matter.

- **Agents' information** $\Omega_t$ is generated by current and past $x$. Agents forecast
  optimally, using the conditional-expectation operator $P[\,\cdot \mid \Omega_t]$.
- **The econometrician's information** $\Sigma_t$ is the closed linear span of current and past
  values of the observed process $y = \{y_t\}$, consisting of the first $q$ components of $x$. The
  econometrician uses *linear least squares projections* onto $\Sigma_t$, which in general have
  larger forecast-error variance than $P[\,\cdot \mid \Omega_t]$.

Because $y$ is linearly regular and full rank, it has a fundamental Wold representation
(see {doc}`13_representation_theory` and {doc}`17_wold_ma`)

$$
y_t = \sum_{j=0}^{\infty} \alpha_j\, v_{t-j}, \qquad E v_t = 0,\quad E v_t v_t' = I,
$$

where $v$ is the process of one-step-ahead innovations that are fundamental for $y$ and span
$\Sigma_t$.

**The economic model.** Partition $y_t = [\,y_{1t}'\ \ y_{2t}'\,]'$ with $y_{1t}$ of dimension
$r$ and $y_{2t}$ of dimension $q-r$. The model is the single (block of) orthogonality condition

```{math}
:label: eq-elre-model
P\big\{\,[A(L)\ ;\ B(L^{-1})]\, y_t \,\big|\, \Omega_{t-\ell}\big\} = 0,
```

where the bracket denotes $A(L)\,y_{1t} + B(L^{-1})\,y_{2t}$. Here $A(L)$ acts on current and
past $y_1$, while $B(L^{-1})$ acts on current and *future* $y_2$, so {eq}`eq-elre-model` equates
to zero the time-$(t-\ell)$ forecast of a combination of past $y_1$ and expected future $y_2$.
The operators are ratios of matrix polynomials,

```{math}
:label: eq-elre-AB
A(z) = A_n(z)/A_d(z), \qquad B(z) = B_n(z)/B_d(z),
```

with $A_n$ an $(r\times r)$ matrix polynomial whose determinant has zeros outside the unit
circle, $B_n$ an $[r\times(q-r)]$ matrix polynomial, and $A_d, B_d$ scalar polynomials with zeros
outside the unit circle.

**Solving the model.** Let $\Lambda_t$ (with $\Sigma_t \subseteq \Lambda_t \subseteq \Omega_t$)
be generated by current and past values of a white noise $w$, $E w_t w_t' = I$. Look for a
time-invariant solution

```{math}
:label: eq-elre-C
y_t = C(L)\, w_t = \begin{bmatrix} C_1(L) \\ C_2(L) \end{bmatrix} w_t,
\qquad C(z) = \sum_{j=0}^{\infty} c_j z^j,\quad \sum_{j=0}^\infty \operatorname{trace}(c_j c_j') < \infty,
```

with $C_1$ of size $(r\times q)$ and $C_2$ of size $[(q-r)\times q]$. Applying the law of iterated
projections (as in {doc}`21_chain_rule`) to {eq}`eq-elre-model` gives
$P\{[A(L)\,;\,B(L^{-1})]y_t \mid \Lambda_t\} = D_1(L)w_t$, where $D_1(z)$ is an $(r\times q)$ matrix
polynomial of degree $\ell-1$ (and $D_1 \equiv 0$ when $\ell = 0$); and {eq}`eq-elre-C` gives
$y_{2t} = D_2(L)w_t$ with $D_2 \equiv C_2$. Matching moving-average coefficients yields

```{math}
:label: eq-elre-solve
A(z)\,C_1(z) + \big[B(z^{-1})\,C_2(z)\big]_+ = D_1(z),
```

so that

```{math}
:label: eq-elre-C1
C_1(z) = A(z)^{-1}\Big\{ D_1(z) - \big[B(z^{-1})\,D_2(z)\big]_+ \Big\},
\qquad C_2(z) = D_2(z).
```

Here $[\,\cdot\,]_+$ is the **annihilation operator** that discards negative powers of $z$:
negative powers multiply future $w$'s, whose projection onto $\Sigma_t$ is zero. This is exactly
the operator in the Wiener–Kolmogorov formula {eq}`eq-62`. In practice one computes
$[B(z^{-1})D_2(z)]_+$ by a partial-fractions/principal-parts calculation — subtracting from
$B(z^{-1})D_2(z)$ the principal parts of its Laurent expansion about the poles inside the unit
disk — which is the residue calculation of {doc}`18a_partial_fractions`.

Equation {eq}`eq-elre-C1` maps a choice of $D(z) = [\,D_1(z)'\ \ D_2(z)'\,]'$ into a solution
$C(z)$; alternative solutions are indexed by alternative admissible $D(z)$'s (those with $D_1$ a
polynomial of degree $\ell-1$ and $D_2$ square-summable). An equivalent and often more useful
characterization is that $C(z)$ solves

```{math}
:label: eq-elre-H
[A(z)\ ;\ B(z^{-1})]\,C(z) = z^{\ell} H(z^{-1})
```

for some $H$ satisfying

```{prf:criterion} Restriction R1
:label: asm-elre-r1

$H(z)$ is an $(r\times q)$ matrix function, analytic on an open set containing
$\{z : |z| \le 1\}$, with $H(0) = 0$.
```

This restatement (with $H(z^{-1}) = z^{-\ell}[D_1(z)+G(z^{-1})]$) is what the identification
analysis of the next part exploits.

## Examples

The general model accommodates a wide range of applications. Three illustrate the pattern; in
each case one simply reads off $A$, $B$, and $\ell$, and then {eq}`eq-elre-C1` delivers the
cross-equation restrictions.

### Example 1: a lognormal model of bond pricing

In the intertemporal asset-pricing model of LeRoy, Rubinstein, Lucas, Breeden, and Cox–Ingersoll–Ross,
the price of an $n$-period pure discount bond satisfies

```{math}
:label: eq-elre-bond
\exp(p_t^n) = P\big[\exp(m_{t+n} - m_t)\mid \Omega_t\big],
```

where $\exp(p_t^n)$ is the bond price and $\exp(m_t)$ is the indirect marginal utility of money.
If $m_t - m_{t-1}$ is a component of a stationary Gaussian $x$, then, up to a constant $c_n$ equal
to half the conditional variance of $m_{t+n}-m_t$,

```{math}
:label: eq-elre-bond2
p_t^n = P\big[(m_{t+1}-m_t) + (m_{t+2}-m_{t+1}) + \cdots + (m_{t+n}-m_{t+n-1}) \mid \Omega_t\big] + c_n .
```

Abstracting from the constant, this is the general model with $y_{1t} = p_t^n$, the first entry
of $y_{2t}$ equal to $m_t - m_{t-1}$, $A(z) = 1$, $B(z) = (z + z^2 + \cdots + z^n)[1\,;\,0]$, and
$\ell = 0$. Hence

```{math}
:label: eq-elre-bond3
C_1(z) = -[1\,;\,0]\,\big[(z^{-1} + z^{-2} + \cdots + z^{-n})\, D_2(z)\big]_+, \qquad C_2(z) = D_2(z).
```

If instead the marginal utility of money is unobserved but the one-period bond price $p_t^1$ is
observed, the law of iterated projections turns {eq}`eq-elre-bond2` into a relation between $p_t^n$
and the expected future path of $p^1$,

$$
p_t^n = P\big[(p_t^1 + p_{t+1}^1 + \cdots + p_{t+n-1}^1)\mid \Omega_t\big] + c_n - n c_1,
$$

now with $B(z) = (1 + z + \cdots + z^{n-1})[1\,;\,0]$. Replacing prices by one-period holding
returns recovers Sargent's (1979) rational expectations model of the term structure.

### Example 2: a present-value model

Let $d_t$ be a dividend or payout and $p_t$ the value of a claim to the dividend stream,

```{math}
:label: eq-elre-pv
p_t = P\Big[\textstyle\sum_{j=0}^{\infty} \lambda^{j}\, d_{t+j} \,\Big|\, \Omega_t\Big],
\qquad 0 < \lambda < 1.
```

This is the general model with $y_{1t} = p_t$, the first entry of $y_{2t}$ equal to $d_t$,
$A(z) = 1$, $B(z) = -\dfrac{1}{1-\lambda z}\,[1\,;\,0]$, and $\ell = 0$. The solution is

```{math}
:label: eq-elre-pv2
C_1(z) = [1\,;\,0]\,\frac{z\,D_2(z) - \lambda\,D_2(\lambda)}{z - \lambda}, \qquad C_2(z) = D_2(z).
```

The present-value form {eq}`eq-elre-pv` is the workhorse behind the stock-price volatility
literature (LeRoy–Porter, Shiller) and government-debt accounting; it is the same geometric
discounting of expected future fundamentals that appears in the Cagan model of
{doc}`22_rational_expectations`, and the "bubble" solutions studied in {doc}`36_bubbles` are
precisely the terms that {eq}`eq-elre-pv2` excludes by requiring a square-summable moving average.

### Example 5: dynamic demand for a factor of production

{cite:t}`sargent1978estimation` and {cite:t}`kennan1979estimation` derive linear factor-demand schedules from quadratic
optimization subject to linear constraints. With one factor and no technology shocks, the demand
function is

```{math}
:label: eq-elre-factor
n_t = \delta\, n_{t-1} - \theta\, P\Big[\textstyle\sum_{j=0}^{\infty} (\beta\delta)^{j}\, p_{t+j} \,\Big|\, \Omega_t\Big],
\qquad 0<\beta<1,\ 0<\delta<1,\ \theta>0,
```

where $n_t$ is the quantity demanded and $p_t$ the factor rental rate. This is the general model
with $y_{1t} = n_t$, the first entry of $y_{2t}$ equal to $p_t$, $A(z) = (1-\delta z)$,
$B(z) = \dfrac{1}{1-\beta\delta z}\,[\theta\,;\,0]$, and $\ell = 0$. Then

```{math}
:label: eq-elre-factor2
C_1(z) = [-\theta\,;\,0]\,\frac{z\,D_2(z) - \beta\delta\, D_2(\beta\delta)}{(1-\delta z)(z - \beta\delta)},
\qquad C_2(z) = D_2(z).
```

The right-hand side of {eq}`eq-elre-factor` is a geometric sum of expected future rental rates —
the same discounted forecast object treated with the chain rule and geometric-lead operator of
{doc}`20_geometric_leads` and {doc}`25_optimal_prediction`. Multiple-factor versions follow the
same template.

## Identification

Take the observed process to have the constrained representation

```{math}
:label: eq-elre-ident
y_t = C(L)\, w_t,
```

with $C(z)$ satisfying {eq}`eq-elre-H`–R1. The population objects the econometrician can hope to
match are the spectral density and autocovariances,

```{math}
:label: eq-elre-spec
S(\theta) = C[e^{-i\theta}]\,C[e^{i\theta}]', \qquad
E(y_t y_{t-\tau}') = \frac{1}{2\pi}\int_{-\pi}^{\pi} e^{i\theta\tau}\, S(\theta)\, d\theta.
```

The difficulty is that $S(\theta)$ does not pin down $C$. Given one **fundamental** moving-average
representation $y_t = F(L)v_t$, every other representation with the same spectral density is
obtained by post-multiplying by an **orthogonal matrix function**,

```{math}
:label: eq-elre-orth
C(z) = F(z)\,U(z), \qquad U[e^{-i\theta}]\,U[e^{i\theta}]' = I \ \text{a.e.},
```

with $U(z) = \sum_{j\ge 0} u_j z^j$ having real, square-summable coefficients. Constant orthogonal
$U$ merely rotate the fundamental innovations; non-constant $U(z)$ move zeros across the unit
circle and generate **non-fundamental** representations, in which $\Lambda_t$ is strictly larger
than $\Sigma_t$. This is the same non-uniqueness of moving-average representations discussed in
{doc}`13_representation_theory` and exploited in {doc}`36a_interpreting_vars`.

The question is how much the cross-equation restrictions narrow this class. Post-multiplying
{eq}`eq-elre-H` by $U(z^{-1})'U(z)$ shows that a rotated $\tilde C(z) = C(z)U(z^{-1})'U(z)$
still satisfies the model if and only if $U$ satisfies

```{prf:criterion} Restriction R2
:label: asm-elre-r2

$H(z)\,U(z)'\,U(z^{-1})$ is analytic on $\{z : |z| < 1\}$.
```

```{prf:lemma}
:label: lem-elre-1

$C(z) = F(z)U(z)$ satisfies the model for some admissible $D(z)$ if and only if $U$ satisfies
{prf:ref}`asm-elre-r2`.
```

A convenient sufficient condition is the stronger

```{prf:criterion} Restriction R3
:label: asm-elre-r3

$U(z)'\,U(z^{-1})$ is analytic on $\{z : |z| < 1\}$.
```

{prf:ref}`asm-elre-r3` holds automatically for constant orthogonal $U$, so fundamental
representations always satisfy the restrictions.

{prf:ref}`asm-elre-r2` is genuinely weaker than {prf:ref}`asm-elre-r3`, so the restrictions do
*not* force fundamentalness. One builds
counterexamples with **Blaschke factors**. For real $\lambda$ with $|\lambda|<1$, set

```{math}
:label: eq-elre-blaschke
\beta(z) = \frac{z - \lambda}{1 - \lambda z}, \qquad \beta(z)\beta(z^{-1}) = 1,
```

and embed $\beta$ in an orthogonal matrix function $U_2(z)$ built from a spectral decomposition of
$H(\lambda)'H(\lambda)$. Because $H(\lambda)$ has a zero column block wherever $U_2(z^{-1})$ has
its pole, the singularity is removable: the resulting $U$ satisfies {prf:ref}`asm-elre-r2` but not
{prf:ref}`asm-elre-r3`, delivering a
non-fundamental, observationally equivalent solution. This is the same Blaschke factorization —
flipping a zero from inside to outside the unit circle without changing the spectral density —
that appears in {doc}`19_signal_extraction` and {doc}`36a_interpreting_vars`. A whole family of
such $U$'s can be manufactured, and finite-order rational parameterizations of $D_2(z)$ do not by
themselves restore identification.

**$D_2(z)$ is generally not identified**, even under prespecified polynomial orders.
The standard remedy is to restrict attention to fundamental representations,

```{math}
:label: eq-elre-fund
C(z) = F(z)\,U \quad \text{for some constant orthogonal } U,
```

but this is ad hoc and leaves the agents' information structure unidentified. Reassuringly, the
non-identification does not invalidate **likelihood-based inference**: the constrained and
unconstrained likelihoods may each have several peaks, but the peaks of interest share the same
value. Moreover, even when $D(z)$ is not identified, the parameters governing $A(z)$ and
$B(z^{-1})$ — the economically interesting ones — often are.

## Restrictions implied for first differences

Covariance stationarity of $y$ is often too strong; a better assumption is that the first
difference of $y_2$ is stationary. The model {eq}`eq-elre-model` transforms neatly under
differencing by a summation-by-parts identity. Writing $B(z) = \sum_{j\ge 0} b_j z^j$, define the
partial sums and the differenced process

```{math}
:label: eq-elre-fd1
b_j^* = \sum_{k=j}^{\infty} b_k, \qquad y_{2t}^* = y_{2t} - y_{2,t-1}, \qquad B^*(z) = \sum_{j=1}^{\infty} b_j^*\, z^j .
```

Since $b_j = b_j^* - b_{j+1}^*$,

```{math}
:label: eq-elre-fd2
B(L^{-1})\, y_{2t} = B^*(L^{-1})\, y_{2t}^* + b_0^*\, y_{2t}.
```

Substituting into {eq}`eq-elre-model` and defining $y_{1t}^* = A(L)y_{1t} + b_0^*\, y_{2t}$ puts
the differenced model back in the general form,

```{math}
:label: eq-elre-fd3
P\big\{\,[\,I\ ;\ B^*(L^{-1})\,]\, y_t^* \,\big|\, \Omega_{t-\ell}\big\} = 0,
\qquad y_t^* = [\,y_{1t}^{*\prime}\ \ y_{2t}^{*\prime}\,]',
```

with $y_t^*$ covariance stationary. This is worth comparing with Sargent (1979), who instead
first-differenced the *whole* relation and projected onto $\Omega_{t-\ell-1}$. Projecting onto the
coarser information set $\Omega_{t-\ell-1}$ discards implications that {eq}`eq-elre-fd3` retains by
projecting onto $\Omega_{t-\ell}$; the representation here therefore imposes more restrictions and
can detect violations that the earlier procedure could not. (The summation-by-parts step is the
same Beveridge–Nelson-style rearrangement that separates a unit-root trend from stationary
dynamics; compare the Wold manipulations of {doc}`17_wold_ma`.)

## Likelihood estimation and inference

To estimate, impose the restrictions on the moving-average coefficients and then maximize a
Gaussian likelihood. Take the pedagogically simplest case: $A(z) = I$,
$B(z) = b_0 + b_1 z + \cdots + b_k z^k$, $\ell = 0$, and a rational parameterization
$D_2(z) = D_n(z)/D_d(z)$ (with $D_d$'s zeros outside the unit circle, conveniently enforced by
Monahan's (1984) parameterization). Then {eq}`eq-elre-C1` collapses to

```{math}
:label: eq-elre-lik1
C_1(z) = -B(z^{-1})\, D_2(z) + G(z^{-1}), \qquad C_2(z) = D_2(z),
```

with $G$ satisfying R1. Clearing $D_d$ gives $C_1(z)D_d(z) = -B(z^{-1})D_n(z) + D_d(z)G(z^{-1})$,
from which $C_1(z) = C_n(z)/D_d(z)$ for a $k$-th order polynomial $C_n$, and $G$ is a polynomial of
order $k$. Collecting powers of $z$ yields

```{math}
:label: eq-elre-lik2
C_n(z) = -B(z^{-1})\, D_n(z) + D_d(z)\, G(z^{-1}),
```

a **recursive linear system**: the equations for the coefficients of $G$ do not involve those of
$C_n$, so one solves a nonsingular triangular system for $G$ and then reads off $C_n$ by simple
matrix multiplication. This representation of the restrictions is both easier to compute and
tighter than the vector-autoregression restrictions Sargent (1979) imposed directly.

With the restricted $C(z)$ in hand, the likelihood can be evaluated two ways.

**Frequency-domain (Whittle) approximation.** Form the finite Fourier transform and periodogram
of the data,

```{math}
:label: eq-elre-lik3
Y(\theta_j) = \sum_{t=1}^{T} y_t\, e^{-i\theta_j t}, \qquad
I(\theta_j) = \tfrac{1}{T}\, Y(\theta_j)\, Y(-\theta_j)', \qquad \theta_j = \tfrac{2\pi j}{T},
```

(omitting $\theta_0$ since means are removed), and approximate the log likelihood by

```{math}
:label: eq-elre-lik4
L = -\frac{nT}{2}\log 2\pi - \frac{1}{2}\sum_{j=1}^{T-1} \log \det S(\theta_j)
   - \frac{1}{2}\sum_{j=1}^{T-1} \operatorname{trace}\!\big[ S(\theta_j)^{-1} I(\theta_j) \big],
```

with $S(\theta_j) = C(e^{-i\theta_j})\,C(e^{i\theta_j})'$ from {eq}`eq-elre-spec`. This is the
spectral (Hannan–Whittle) likelihood, which draws on the spectral apparatus of
{doc}`06_spectrum` and {doc}`07_cross_spectrum`; symmetry across $\theta$ and $2\pi-\theta$ halves
the work. The free parameters of $D(z)$ are estimated by maximizing {eq}`eq-elre-lik4` subject to
the restrictions {eq}`eq-elre-C1`.

**Time-domain filtering.** When $D(z)$ is restricted so that

$$
C(z) = \big[\,C_n(z)/D_d(z)\ ;\ D_n(z)/D_d(z)\,\big]'
$$

is rational, $y$ is a vector ARMA process with a state-space representation, and the *exact*
Gaussian likelihood follows from the conditional-density factorization

$$
f(y_1,\ldots,y_T) = f(y_T \mid y_1,\ldots,y_{T-1})\cdots f(y_2\mid y_1)\, f(y_1),
$$

whose conditional means and covariances are produced recursively by the Kalman filter — the same
innovations/filtering recursions developed in {doc}`26_optimal_filtering` and
{doc}`30_filtering_projections`, initialized by a doubling algorithm for the stationary
covariance. This route avoids the frequency-domain approximation at the cost of restricting
$D(z)$ enough to guarantee a rational $C(z)$.

As with identification, taking $A(z)$ and $B(z^{-1})$ as given and estimating $D(z)$ is the hard
case; when $A$ and $B$ depend on a finite parameter vector, those structural parameters are
frequently identified and estimable even where $D(z)$ is not.

## References

```{bibliography}
:labelprefix: EL
:filter: key in {"hansensargent1980formulating", "hansensargent1991exact", "kennan1979estimation", "monahan1984note", "sargent1978estimation", "sargent1979note", "whittle1953estimation"}
```
