12. Linear Least Squares Prediction#

(a) The Wiener–Kolmogorov Formula#

The Wold moving average representation is useful for representing the linear least squares \(u\)-step ahead prediction for a purely linearly indeterministic process. Let \(x(t)\) be a purely linearly indeterministic process with Wold moving average representation

(62)#\[ x(t) = \int^\infty_0 p(s)\, w(t-s)\, ds \]

where \(\int^\infty_0 p(s)^2\, ds < +\infty\) and \(w(t)\) is a fundamental white noise for \(x(t)\). Recall that the property that \(w(t)\) is a fundamental white noise for \(x(t)\) means equality of the linear spaces

\[\begin{split} \begin{aligned} H_x(-\infty,\, t) \equiv\ & [\, y(t) : y(t) = \int^\infty_0 b(s)\, x(t-s)\, ds, \\ &\ \text{ for any }\ b(s) \in L_2\,[0,\, \infty) \,] \end{aligned} \end{split}\]

and

\[\begin{split} \begin{aligned} H_w(-\infty,\, t) \equiv\ & [\, z(t) : z(t) = \int^\infty_0 h(s)\, w(t-s)\, ds \\ &\ \text{ for any }\ h(s) \in L_2\,[0,\, \infty) \,] \end{aligned} \end{split}\]

where \(L_2\,[0,\, \infty)\) is the space of square integrable functions, i.e., functions \(b(s)\) for \(0 \leq s < \infty\) such that \(\int^\infty_0 b(s)^2\, ds < +\infty\). The equality of these spaces means that lagged \(x\)’s contain the same amount of information as lagged \(w\)’s.

Since (62) holds for all \(t\), we have

\[ x(t + u) = \int^\infty_{s=-u} p(s+u)\, w(t-s)\, ds,\ \text{ for }\ u \geq 0. \]

Using the identity of the linear spaces \(H_x(-\infty,\,t)\) and \(H_w(-\infty,\,t)\), we have that

(63)#\[ E\, [\, x(t+u) \mid x(v),\, v \leq t \,] = \int^\infty_{s=0} p(s+u)\, w(t-s)\, ds. \]

Equation (63) is the continuous time Wiener–Kolmogorov formula. It rests on the equality of the spaces spanned by past \(x\)’s and past fundamental innovations \(w\)’s established in Wold’s theorem (Theorem 10 of 8. Spectral Densities); 15. State-Space Models, the Kalman Filter, and Spectral Factorization gives the equivalent state-space form, in which \(w\) becomes the Kalman innovations process and this same forecast is computed recursively.

Using operational calculus, the formula can be expressed as

(64)#\[ E_t\, x(t+u) = [\, \tilde P(D)\, e^{Du} \,]_+\, w(t) \]

where \([\, \tilde P(s)\, e^{su} \,]_+\) is the time function formed by taking the inverse Laplace transform of \(\tilde P(s)\, e^{su}\), and then multiplying it by the Heaviside unit step function (i.e., setting values of the time function for \(t < 0\) equal to zero, while leaving values of the function for \(t \geq 0\) unaltered). The operator \([\,\cdot\,]_+\) is known as the annihilation operator. Note that \(e^{uD}\) is the operator that shifts a time function ahead \(u\) units. Reading property 4 (Delay) of Table 2 in the direction of an advance rather than a delay, \(e^{su}\, \tilde P(s)\) is the Laplace transform of the advanced function \(p(t+u)\); annihilating its negative part then leaves the kernel \(p(s+u),\ s \geq 0\), of (63). This \(e^{uD}\) convention is the one used throughout the book, in 19. Prediction Formulas for Continuous Time Linear Rational Expectations Models and in 20. Aggregation Over Time and the Inverse Optimal Predictor Problem for Adaptive Expectations in Continuous Time.

As a check, for \(\tilde P(s) = 1/(a+s)\) the inverse transform of \(e^{su}/(a+s)\) is \(e^{-a(t+u)}\) for \(t + u \geq 0\); annihilating \(t<0\) and transforming back gives \(e^{-au}/(a+s)\), which is (65).

As an example of the use of formula (63), let \(x(t)\) be governed by the first order stochastic differential equation

\[ (D+a)\, x(t) = w(t),\qquad a > 0 \]

so that \(p(\tau) = e^{-a\tau}\). Then formula (63) gives

\[\begin{split} \begin{aligned} E_t\, x(t+u) &= \int^\infty_0 e^{-a(s+u)}\, w(t-s)\, ds \\ &= e^{-au} \int^\infty_0 e^{-as}\, w(t-s)\, ds = e^{-au}\ \frac{1}{D + a}\, w(t) \end{aligned} \end{split}\]

or

(65)#\[ E_t\, x(t+u) = e^{-au}\, x(t). \]

(b) A Formula for Predicting “Geometric Distributed Leads”#

Such geometric distributed leads are the present values that appear in every asset-pricing equation, permanent-income model, and quadratic-adjustment-cost Euler equation; they are also the continuous-time counterpart of the discounted expected sums that define the optimal feedforward decision rules of 16. Faster Methods for Solving Recursive Linear Models of Dynamic Economies. Evaluating them is the central computational step in solving continuous-time rational expectations models.

In linear rational expectation models, there often appear terms of the form

(66)#\[ E_t \int^\infty_0 e^{\rho u}\, x(t+u)\, du \qquad re(\rho) < 0 \]

where \(x(t)\) is a covariance stationary stochastic process. Where \(x(t)\) is governed by the first order Markov process \((D+a)\, x(t) = w(t)\), equation (65) implies that the linear least squares forecast of the geometric distributed lead (66) is given by

\[ \int^\infty_0 e^{\rho u}\, E_t\, x(t+u)\, du = \left( \int^\infty_0 e^{\rho u}\, e^{-au}\, du \right) x(t) = \frac{1}{a - \rho}\, x(t). \]

An approach to the evaluation of (66) which readily generalizes to \(x(t)\)’s governed by higher order linear differential equations is as follows. Denote the geometric distributed lead to be forecast as

(67)#\[\begin{split} \begin{aligned} x(t)^{\ast} &= \int^\infty_0 e^{\rho u}\, x(t+u)\, du \\ x(t)^{\ast} &= \left( \frac{-1}{\rho + D} \right)\ \left( \frac{1}{a+D} \right)\, w(t) \end{aligned} \end{split}\]

where for \(re(\rho) < 0\), \(-1/(\rho + s)\) is the (two-sided) Laplace transform of the time function equal to \(e^{-\rho u}\) for \(u \leq 0\) and \(0\) for \(u > 0\).

Obtaining a partial fraction representation of the right side of (67) gives

\[ x(t)^{\ast} = \frac{1}{a-\rho}\ \left[ \left( \frac{-1}{\rho + D} \right) + \left( \frac{1}{a+D} \right) \right]\ w(t) \]

or

\[ x(t)^{\ast} = \left( \frac{1}{a-\rho} \right)\ \left[ -\int^\infty_0 e^{\rho s}\, w(t+s)\, ds + \int^\infty_0 e^{-as}\, w(t-s)\, ds \right] \]

It then follows that

(68)#\[\begin{split} \begin{aligned} E_t\, x(t)^{\ast} &= \left( \frac{1}{a-\rho} \right) \int^\infty_0 e^{-as}\, w(t-s)\, ds = \left( \frac{1}{a-\rho} \right)\ \frac{1}{a + D}\ w(t) \\ &= \left( \frac{1}{a - \rho} \right) x(t). \end{aligned} \end{split}\]

This approach generalizes readily as follows. Represent equation (68) as

(69)#\[\begin{split} \begin{aligned} E_t\, x(t)^{\ast} &= \left( \frac{1}{a - \rho} \right)\ \left( \frac{1}{a+D} \right)\, w(t) \\ E_t\, x(t)^{\ast} &= \left[ \frac{-\tilde P(D) + \tilde P(-\rho)}{D+\rho} \right]\, w(t) \end{aligned} \end{split}\]

where \(\tilde P(D) = 1/(a+D)\). As it happens, Equation (69) holds for any \(\tilde P(D)\), where \(\tilde P(s)\) is the Laplace transform of a squared summable function \(p(\tau)\) concentrated on \(\tau \in [0,\, \infty)\). Thus, where

\[ x(t) = \int^\infty_0 p(\tau)\, w(t-\tau)\, d\tau \]

we claim that the generalization of (69) is

(70)#\[ E_t \int^\infty_0 e^{\rho s}\, x(t+s)\, ds = \left[ \frac{-\tilde P(D) + \tilde P(-\rho)}{D+\rho} \right]\, w(t). \]

The general formula (70), valid for any rational \(\tilde P(D)\), is established in 19. Prediction Formulas for Continuous Time Linear Rational Expectations Models.

Exercises#

There is a direct route to the kernel of a geometric distributed lead that provides an independent check on (70). Writing \(x(t+s) = \int_0^\infty p(\tau)w(t+s-\tau)d\tau\) and discarding the terms dated after \(t\) leaves \(E_t\, x(t+s) = \int_0^\infty p(\tau+s)\, w(t-\tau)\, d\tau\), which is (63). Multiplying by \(e^{\rho s}\) and integrating,

(71)#\[E_t \int_0^\infty e^{\rho s} x(t+s)\, ds = \int_0^\infty g(\tau)\, w(t-\tau)\, d\tau, \qquad g(\tau) = \int_0^\infty e^{\rho u}\, p(\tau + u)\, du .\]

Formula (71) computes the same object as (70), but in the time domain and by quadrature rather than by partial fractions.

import numpy as np
from scipy.integrate import quad

Exercise 14

Take the second-order process \((D+a_1)(D+a_2)x(t) = w(t)\) with \(a_1 = 0.6\), \(a_2 = 1.7\), whose kernel is \(p(\tau) = (e^{-a_1\tau} - e^{-a_2\tau})/(a_2 - a_1)\), and let \(\rho = -0.4\).

(a) Show by partial fractions that (70) gives

\[ g(\tau) = \frac{1}{a_2 - a_1}\left[\frac{e^{-a_1\tau}}{a_1 - \rho} - \frac{e^{-a_2\tau}}{a_2 - \rho}\right], \]

and verify this closed form against the quadrature (71).

(b) Confirm that \(g(0) = \tilde P(-\rho) = 1/[(a_1-\rho)(a_2-\rho)]\), as the initial value theorem of 9. Characterizations of Mean Square Differentiability and Mean Square Continuity requires. Note that \(g(0) \neq 0\) even though \(p(0) = 0\). 13. Locally Unpredictable Stochastic Processes exploits that fact.

(c) Recover the first-order case by letting \(a_2 \to \infty\). Note that the kernel itself vanishes in that limit, so the process must be rescaled: \(a_2\, p(\tau) \to e^{-a_1\tau}\), and correspondingly \(a_2\, g(0) \to 1/(a_1 - \rho)\), which is the coefficient in (68).