Exact Linear Rational Expectations Models#
Note
This section is based on Lars Peter Hansen and Thomas J. Sargent, “Exact Linear Rational Expectations Models: Specification and Estimation,” Chapter 3 of Rational Expectations Econometrics (Westview Press, 1991). We follow the paper’s Sections 1–5, present only Examples 1, 2, and 5, and have shortened and reorganized the exposition. The technical identification arguments of Section 3 are summarized rather than reproduced in full.
This section follows A Difficulty in Interpreting Vector Autoregressions, reversing the order of the two chapters in the source. Some Applications to Rational Expectations Models and Sims’s Application to Money and Income raise the difficulty that Chapter 4 diagnoses, so that chapter comes first here; the constructive apparatus of Chapter 3 then answers it.
A distinguishing feature of econometric models that build in rational expectations is a set of cross-equation restrictions: the parameters of the equations describing variables that people choose are inherited from the equations describing the stochastic environment people forecast. Because decisions depend on forecasts, and forecasts depend on the laws of motion of the environment, the decision equations cannot have free parameters. Even for models that are linear in the variables, these restrictions are typically nonlinear in the underlying parameters.
This section describes a convenient way to characterize and impose those restrictions for a class of exact linear rational expectations models — models in which an exact linear relation ties forecasts of future values of one set of variables to current and past values of another set, and every variable entering the relation is observed by the econometrician. That last requirement is what makes the model “exact”: the relation the econometrician confronts would hold without error if agents could forecast perfectly. Many models of the term structure, of stock prices, of consumption and permanent income, and of the dynamic demand for factors of production belong to this class.
The strategy throughout is to work with a vector moving-average representation of the observed process and to deduce the restrictions by straightforward applications of the Wiener–Kolmogorov prediction formula (248) — the same annihilation-operator machinery used in Signal Extraction Problems and The Residue Theorem Behind Partial Fractions. Once the restrictions are in hand, the constrained moving average is nested inside a less constrained one, and the model can be estimated and tested by maximum likelihood.
The general model#
Let \(x = \{x_t\}\) be a \(p\)-dimensional, covariance-stationary process with mean zero. Two information sets matter.
Agents’ information \(\Omega_t\) is generated by current and past \(x\). Agents forecast optimally, using the conditional-expectation operator \(P[\,\cdot \mid \Omega_t]\).
The econometrician’s information \(\Sigma_t\) is the closed linear span of current and past values of the observed process \(y = \{y_t\}\), consisting of the first \(q\) components of \(x\). The econometrician uses linear least squares projections onto \(\Sigma_t\), which in general have larger forecast-error variance than \(P[\,\cdot \mid \Omega_t]\).
Because \(y\) is linearly regular and full rank, it has a fundamental Wold representation (see Representation Theory and Finding a Wold Representation: mth Order Moving Average)
where \(v\) is the process of one-step-ahead innovations that are fundamental for \(y\) and span \(\Sigma_t\).
The economic model. Partition \(y_t = [\,y_{1t}'\ \ y_{2t}'\,]'\) with \(y_{1t}\) of dimension \(r\) and \(y_{2t}\) of dimension \(q-r\). The model is the single (block of) orthogonality condition
where the bracket denotes \(A(L)\,y_{1t} + B(L^{-1})\,y_{2t}\). Here \(A(L)\) acts on current and past \(y_1\), while \(B(L^{-1})\) acts on current and future \(y_2\), so (463) equates to zero the time-\((t-\ell)\) forecast of a combination of past \(y_1\) and expected future \(y_2\). The operators are ratios of matrix polynomials,
with \(A_n\) an \((r\times r)\) matrix polynomial whose determinant has zeros outside the unit circle, \(B_n\) an \([r\times(q-r)]\) matrix polynomial, and \(A_d, B_d\) scalar polynomials with zeros outside the unit circle.
Solving the model. Let \(\Lambda_t\) (with \(\Sigma_t \subseteq \Lambda_t \subseteq \Omega_t\)) be generated by current and past values of a white noise \(w\), \(E w_t w_t' = I\). Look for a time-invariant solution
with \(C_1\) of size \((r\times q)\) and \(C_2\) of size \([(q-r)\times q]\). Applying the law of iterated projections (as in The Chain Rule of Forecasting) to (463) gives \(P\{[A(L)\,;\,B(L^{-1})]y_t \mid \Lambda_t\} = D_1(L)w_t\), where \(D_1(z)\) is an \((r\times q)\) matrix polynomial of degree \(\ell-1\) (and \(D_1 \equiv 0\) when \(\ell = 0\)); and (465) gives \(y_{2t} = D_2(L)w_t\) with \(D_2 \equiv C_2\). Matching moving-average coefficients yields
so that
Here \([\,\cdot\,]_+\) is the annihilation operator that discards negative powers of \(z\): negative powers multiply future \(w\)’s, whose projection onto \(\Sigma_t\) is zero. This is exactly the operator in the Wiener–Kolmogorov formula (248). In practice one computes \([B(z^{-1})D_2(z)]_+\) by a partial-fractions/principal-parts calculation — subtracting from \(B(z^{-1})D_2(z)\) the principal parts of its Laurent expansion about the poles inside the unit disk — which is the residue calculation of The Residue Theorem Behind Partial Fractions.
Equation (467) maps a choice of \(D(z) = [\,D_1(z)'\ \ D_2(z)'\,]'\) into a solution \(C(z)\); alternative solutions are indexed by alternative admissible \(D(z)\)’s (those with \(D_1\) a polynomial of degree \(\ell-1\) and \(D_2\) square-summable). An equivalent and often more useful characterization is that \(C(z)\) solves
for some \(H\) satisfying
Criterion 1 (Restriction R1)
\(H(z)\) is an \((r\times q)\) matrix function, analytic on an open set containing \(\{z : |z| \le 1\}\), with \(H(0) = 0\).
This restatement (with \(H(z^{-1}) = z^{-\ell}[D_1(z)+G(z^{-1})]\)) is what the identification analysis of the next part exploits.
Examples#
The general model accommodates a wide range of applications. Three illustrate the pattern; in each case one simply reads off \(A\), \(B\), and \(\ell\), and then (467) delivers the cross-equation restrictions.
Example 1: a lognormal model of bond pricing#
In the intertemporal asset-pricing model of LeRoy, Rubinstein, Lucas, Breeden, and Cox–Ingersoll–Ross, the price of an \(n\)-period pure discount bond satisfies
where \(\exp(p_t^n)\) is the bond price and \(\exp(m_t)\) is the indirect marginal utility of money. If \(m_t - m_{t-1}\) is a component of a stationary Gaussian \(x\), then, up to a constant \(c_n\) equal to half the conditional variance of \(m_{t+n}-m_t\),
Abstracting from the constant, this is the general model with \(y_{1t} = p_t^n\), the first entry of \(y_{2t}\) equal to \(m_t - m_{t-1}\), \(A(z) = 1\), \(B(z) = (z + z^2 + \cdots + z^n)[1\,;\,0]\), and \(\ell = 0\). Hence
If instead the marginal utility of money is unobserved but the one-period bond price \(p_t^1\) is observed, the law of iterated projections turns (470) into a relation between \(p_t^n\) and the expected future path of \(p^1\),
now with \(B(z) = (1 + z + \cdots + z^{n-1})[1\,;\,0]\). Replacing prices by one-period holding returns recovers Sargent’s (1979) rational expectations model of the term structure.
Example 2: a present-value model#
Let \(d_t\) be a dividend or payout and \(p_t\) the value of a claim to the dividend stream,
This is the general model with \(y_{1t} = p_t\), the first entry of \(y_{2t}\) equal to \(d_t\), \(A(z) = 1\), \(B(z) = -\dfrac{1}{1-\lambda z}\,[1\,;\,0]\), and \(\ell = 0\). The solution is
The present-value form (472) is the workhorse behind the stock-price volatility literature (LeRoy–Porter, Shiller) and government-debt accounting; it is the same geometric discounting of expected future fundamentals that appears in the Cagan model of Some Applications to Rational Expectations Models, and the “bubble” solutions studied in Bubbles are precisely the terms that (473) excludes by requiring a square-summable moving average.
Example 5: dynamic demand for a factor of production#
Sargent (1978) and Kennan (1979) derive linear factor-demand schedules from quadratic optimization subject to linear constraints. With one factor and no technology shocks, the demand function is
where \(n_t\) is the quantity demanded and \(p_t\) the factor rental rate. This is the general model with \(y_{1t} = n_t\), the first entry of \(y_{2t}\) equal to \(p_t\), \(A(z) = (1-\delta z)\), \(B(z) = \dfrac{1}{1-\beta\delta z}\,[\theta\,;\,0]\), and \(\ell = 0\). Then
The right-hand side of (474) is a geometric sum of expected future rental rates — the same discounted forecast object treated with the chain rule and geometric-lead operator of Predicting Geometric Distributed Leads and Optimal Prediction: Compact Notation. Multiple-factor versions follow the same template.
Identification#
Take the observed process to have the constrained representation
with \(C(z)\) satisfying (468)–R1. The population objects the econometrician can hope to match are the spectral density and autocovariances,
The difficulty is that \(S(\theta)\) does not pin down \(C\). Given one fundamental moving-average representation \(y_t = F(L)v_t\), every other representation with the same spectral density is obtained by post-multiplying by an orthogonal matrix function,
with \(U(z) = \sum_{j\ge 0} u_j z^j\) having real, square-summable coefficients. Constant orthogonal \(U\) merely rotate the fundamental innovations; non-constant \(U(z)\) move zeros across the unit circle and generate non-fundamental representations, in which \(\Lambda_t\) is strictly larger than \(\Sigma_t\). This is the same non-uniqueness of moving-average representations discussed in Representation Theory and exploited in A Difficulty in Interpreting Vector Autoregressions.
The question is how much the cross-equation restrictions narrow this class. Post-multiplying (468) by \(U(z^{-1})'U(z)\) shows that a rotated \(\tilde C(z) = C(z)U(z^{-1})'U(z)\) still satisfies the model if and only if \(U\) satisfies
Criterion 2 (Restriction R2)
\(H(z)\,U(z)'\,U(z^{-1})\) is analytic on \(\{z : |z| < 1\}\).
Lemma 1
\(C(z) = F(z)U(z)\) satisfies the model for some admissible \(D(z)\) if and only if \(U\) satisfies Criterion 2.
A convenient sufficient condition is the stronger
Criterion 3 (Restriction R3)
\(U(z)'\,U(z^{-1})\) is analytic on \(\{z : |z| < 1\}\).
Criterion 3 holds automatically for constant orthogonal \(U\), so fundamental representations always satisfy the restrictions.
Criterion 2 is genuinely weaker than Criterion 3, so the restrictions do not force fundamentalness. One builds counterexamples with Blaschke factors. For real \(\lambda\) with \(|\lambda|<1\), set
and embed \(\beta\) in an orthogonal matrix function \(U_2(z)\) built from a spectral decomposition of \(H(\lambda)'H(\lambda)\). Because \(H(\lambda)\) has a zero column block wherever \(U_2(z^{-1})\) has its pole, the singularity is removable: the resulting \(U\) satisfies Criterion 2 but not Criterion 3, delivering a non-fundamental, observationally equivalent solution. This is the same Blaschke factorization — flipping a zero from inside to outside the unit circle without changing the spectral density — that appears in Signal Extraction Problems and A Difficulty in Interpreting Vector Autoregressions. A whole family of such \(U\)’s can be manufactured, and finite-order rational parameterizations of \(D_2(z)\) do not by themselves restore identification.
\(D_2(z)\) is generally not identified, even under prespecified polynomial orders. The standard remedy is to restrict attention to fundamental representations,
but this is ad hoc and leaves the agents’ information structure unidentified. Reassuringly, the non-identification does not invalidate likelihood-based inference: the constrained and unconstrained likelihoods may each have several peaks, but the peaks of interest share the same value. Moreover, even when \(D(z)\) is not identified, the parameters governing \(A(z)\) and \(B(z^{-1})\) — the economically interesting ones — often are.
Restrictions implied for first differences#
Covariance stationarity of \(y\) is often too strong; a better assumption is that the first difference of \(y_2\) is stationary. The model (463) transforms neatly under differencing by a summation-by-parts identity. Writing \(B(z) = \sum_{j\ge 0} b_j z^j\), define the partial sums and the differenced process
Since \(b_j = b_j^* - b_{j+1}^*\),
Substituting into (463) and defining \(y_{1t}^* = A(L)y_{1t} + b_0^*\, y_{2t}\) puts the differenced model back in the general form,
with \(y_t^*\) covariance stationary. This is worth comparing with Sargent (1979), who instead first-differenced the whole relation and projected onto \(\Omega_{t-\ell-1}\). Projecting onto the coarser information set \(\Omega_{t-\ell-1}\) discards implications that (483) retains by projecting onto \(\Omega_{t-\ell}\); the representation here therefore imposes more restrictions and can detect violations that the earlier procedure could not. (The summation-by-parts step is the same Beveridge–Nelson-style rearrangement that separates a unit-root trend from stationary dynamics; compare the Wold manipulations of Finding a Wold Representation: mth Order Moving Average.)
Likelihood estimation and inference#
To estimate, impose the restrictions on the moving-average coefficients and then maximize a Gaussian likelihood. Take the pedagogically simplest case: \(A(z) = I\), \(B(z) = b_0 + b_1 z + \cdots + b_k z^k\), \(\ell = 0\), and a rational parameterization \(D_2(z) = D_n(z)/D_d(z)\) (with \(D_d\)’s zeros outside the unit circle, conveniently enforced by Monahan’s (1984) parameterization). Then (467) collapses to
with \(G\) satisfying R1. Clearing \(D_d\) gives \(C_1(z)D_d(z) = -B(z^{-1})D_n(z) + D_d(z)G(z^{-1})\), from which \(C_1(z) = C_n(z)/D_d(z)\) for a \(k\)-th order polynomial \(C_n\), and \(G\) is a polynomial of order \(k\). Collecting powers of \(z\) yields
a recursive linear system: the equations for the coefficients of \(G\) do not involve those of \(C_n\), so one solves a nonsingular triangular system for \(G\) and then reads off \(C_n\) by simple matrix multiplication. This representation of the restrictions is both easier to compute and tighter than the vector-autoregression restrictions Sargent (1979) imposed directly.
With the restricted \(C(z)\) in hand, the likelihood can be evaluated two ways.
Frequency-domain (Whittle) approximation. Form the finite Fourier transform and periodogram of the data,
(omitting \(\theta_0\) since means are removed), and approximate the log likelihood by
with \(S(\theta_j) = C(e^{-i\theta_j})\,C(e^{i\theta_j})'\) from (477). This is the spectral (Hannan–Whittle) likelihood, which draws on the spectral apparatus of The Spectrum and The Cross Spectrum; symmetry across \(\theta\) and \(2\pi-\theta\) halves the work. The free parameters of \(D(z)\) are estimated by maximizing (487) subject to the restrictions (467).
Time-domain filtering. When \(D(z)\) is restricted so that
is rational, \(y\) is a vector ARMA process with a state-space representation, and the exact Gaussian likelihood follows from the conditional-density factorization
whose conditional means and covariances are produced recursively by the Kalman filter — the same innovations/filtering recursions developed in Optimal Filtering Formula and The Effects of Filtering on One-Sided and Two-Sided Projections, initialized by a doubling algorithm for the stationary covariance. This route avoids the frequency-domain approximation at the cost of restricting \(D(z)\) enough to guarantee a rational \(C(z)\).
As with identification, taking \(A(z)\) and \(B(z^{-1})\) as given and estimating \(D(z)\) is the hard case; when \(A\) and \(B\) depend on a finite parameter vector, those structural parameters are frequently identified and estimable even where \(D(z)\) is not.
References#
Lars Peter Hansen and Thomas J. Sargent. Formulating and estimating dynamic linear rational expectations models. Journal of Economic Dynamics and Control, 2(1):7–46, 1980.
Lars Peter Hansen and Thomas J. Sargent. Exact linear rational expectations models: specification and estimation. In Rational Expectations Econometrics, chapter 3. Westview Press, Boulder, CO, 1991.
John Kennan. The estimation of partial adjustment models with rational expectations. Econometrica, 47(6):1441–1455, 1979.
John F. Monahan. A note on enforcing stationarity in autoregressive-moving average models. Biometrika, 71(2):403–404, 1984.
Thomas J. Sargent. Estimation of dynamic labor demand schedules under rational expectations. Journal of Political Economy, 86(6):1009–1044, 1978.
Thomas J. Sargent. A note on maximum likelihood estimation of the rational expectations model of the term structure. Journal of Monetary Economics, 5(1):133–143, 1979.
Peter Whittle. Estimation and information in stationary time series. Arkiv för Matematik, 2(5):423–434, 1953.