Money Demand in Hyperinflations: A Misspecified Regression and Sims’s Approximation Formula#
Note
This section is based on Thomas J. Sargent, “The Demand for Money During Hyperinflations under Rational Expectations: I,” International Economic Review 18(1), 59–82 (1977). We follow the paper’s projection theory and summarize its quantitative findings; we omit the details of the maximum likelihood estimation. The full computational treatment, with Python code for the simulations, the estimator, and the empirical tables, is the QuantEcon lecture Demand for Money during Hyperinflations under Rational Expectations.
This section draws on four earlier ones. The rational expectations Cagan model supplies the equilibrium. The Granger causality and econometric exogeneity theorem and its money–income application supply the exogeneity argument. The phase and leading of The Cross Spectrum and A Digression on Leading Indicators supply the reading of Cagan’s paradox. Sims’s approximation error formula (385), stated in Seasonality and Approximation Errors, organizes the whole.
Cagan’s paradox#
Cagan’s (1956) classic study of seven hyperinflations fit a demand schedule for real balances,
where \(m_t\) is log money, \(p_t\) log price level, \(\pi_t\) the public’s expected inflation, and \(u_t\) a mean-zero disturbance. Cagan estimated \(\alpha\) — the semi-elasticity of real balances with respect to expected inflation — and used it to compute the inflation rate \(-1/\alpha\) that would maximize the inflation-tax revenue a money-printing government can extract. For every one of the seven hyperinflations, the reciprocal of Cagan’s estimate of \(-\alpha\) came out far below the average actual inflation rate: money creators appeared to be inflating at rates wildly in excess of the revenue-maximizing rate. The natural suspicion is that this paradox is a statistical artifact — a consequence of a biased estimate of \(\alpha\). This section shows that it is, and that the bias is an instance of Sims’s approximation-error formula applied to a misspecified regression.
Cagan’s model under rational expectations#
Cagan assumed adaptive expectations, \(\pi_t=\dfrac{1-\lambda}{1-\lambda L}\,x_t\), where \(x_t\equiv p_t-p_{t-1}\) is inflation and \(L\) the lag operator. Sargent and Wallace (1973) observed that under rational expectations, \(\pi_t = E_t x_{t+1}\), and solving the resulting forward difference equation with \(\lvert{-\alpha}/(1-\alpha)\rvert<1\) gives
with \(\mu_t\equiv m_t-m_{t-1}\) the money-growth rate. Equation (489) is a geometric distributed lead of the same kind studied throughout this chapter: expected inflation is a discounted sum of expected future money growth, so — with the stochastic process for money creation held fixed — “money causes inflation.” Adaptive expectations coincide with rational expectations only under restrictions on \(u\) and \(\mu\); Sargent studies two sufficient conditions: \(u_t\) is a random walk, \(u_t=u_{t-1}+\eta_t\) (so \(E_t(u_{t+j}-u_{t+j-1})=0\)), and the money-creation rule
with \(\varepsilon_t,\eta_t\) serially uncorrelated, mean zero. Under (490) the geometric sum in (489) collapses to \(\pi_t=\frac{1-\lambda}{1-\lambda L}x_t\), so adaptive expectations are rational. Rule (490) describes a government that prints money in response to inflation — a “real-bills” regime of the sort German officials repeatedly invoked to argue that money was responding to inflation rather than causing it.
Under these conditions inflation and money growth form a bivariate system that reduces, after first-differencing, to a first-order moving average driven by \((\varepsilon_t,\eta_t)\). Writing \(\phi\equiv(\lambda+\alpha(1-\lambda))^{-1}\),
The bivariate Wold representation and Granger causality#
To find what Cagan’s regression converges to we need the fundamental (Wold) representation of \((\Delta x_t,\Delta\mu_t)\). Project the money-supply shock on the inflation shock: write \(\varepsilon_t=\rho(\varepsilon_t-\eta_t)+v_t\) with
so that \(v_t\perp(\varepsilon_t-\eta_t)\). Substituting into (491) gives the triangular bivariate Wold representation
with jointly fundamental white noises \((\varepsilon_t-\eta_t,\,v_t)\). The lower-triangular structure — \(\Delta x\) loads only on the first shock, \(\Delta\mu\) on both — is exactly the condition of Sims's theorem: \(\Delta x\) is econometrically exogenous with respect to \(\Delta\mu\), so inflation Granger-causes money creation but not conversely. This is the same one-sidedness that organizes Sims’s money–income test, here running from prices to money rather than money to income.
A caution from A Digression on Leading Indicators applies with full force. That \(x\) causes \(\mu\) in Granger’s sense does not mean \(x\) leads \(\mu\) in any National-Bureau sense. Indeed, the equilibrium implies \(\mu_t = x_t + a_{2t}-a_{1t}\) (with \(a_{1t},a_{2t}\) the innovations in \(x,\mu\)), so \(x_t\) and \(\mu_t\) are in phase — their cross-spectrum has zero phase at every frequency (The Cross Spectrum). Evidence that inflation leads money would be evidence against the model. Granger causality and phase leads are different things, exactly as A Digression on Leading Indicators warned.
Finally, the Granger-causal pattern here is a property of the particular money rule (490), not an invariant feature of the economy: change the money-supply regime and the causal ordering can change. This is the distinction — central to Chapter XIV and to Exact Linear Rational Expectations Models — between Granger causality and invariance under intervention.
Cagan’s regression as a misspecified distributed lag#
Under the two sufficient conditions the structural money-demand schedule (488) becomes a one-parameter distributed-lag regression of real balances on inflation,
Cagan read the least-squares fit of (494) as delivering the structural \(\alpha\). But this regression is misspecified: the disturbance \(u_t\) is not orthogonal to the inflation process. Under the equilibrium money rule (490), money growth responds to inflation, so \(\pi_t\) (the regressor) is correlated with \(u_t\) (the error) — a simultaneity that biases ordinary least squares. Computing the actual population projection of \(m_t-p_t\) on the whole inflation process — substitute the Wold representation (493) and apply the summation operator \((1-L)^{-1}\) — gives
with \(\bar u_t\) a random walk orthogonal to the \(x\) process. Comparing (495) with the structural form (494):
Cagan’s estimator of the shape parameter \(\lambda\) is consistent;
Cagan’s estimator of the slope \(\alpha\) is not, converging instead to
The probability limit is a weighted average of the structural \(\alpha\) and the Wallace–Sargent value \(-\lambda/(1-\lambda)\), with weight \(\rho\). Two polar cases pin down the intuition: if there is no noise in portfolio balance (\(\eta_t\equiv0\), so \(\rho=1\)), OLS is consistent, \(\operatorname{plim}\hat\alpha=\alpha\); if \(\rho=0\), the estimator collapses to the pure Wallace–Sargent value \(-\lambda/(1-\lambda)\), independent of the true \(\alpha\). When the true \(\alpha\) is more negative than \(-\lambda/(1-\lambda)\) — the empirically relevant case — the bias pulls \(\hat\alpha\) toward zero, shrinking \(-1/\hat\alpha\) below the true revenue-maximizing rate. That is Cagan’s paradox, manufactured by misspecification. (Consistent with the random-walk residual \(\bar u_t\) in (495), both Cagan and Barro reported highly serially correlated residuals and very low Durbin–Watson statistics.)
Sims’s approximation-error formula#
Equation (496) is an instance of Sims’s frequency-domain approximation-error formula. Recall (385) from Seasonality and Approximation Errors, derived in Exercise 33: if the true projection of \(y_t\) on a covariance-stationary process \(x_t\) has transfer function \(b^0(L)\), and an econometrician fits, by least squares, a constrained distributed lag \(b^1(L)\) drawn from a restricted class, then in population least squares chooses \(b^1\) to minimize the spectral-density-weighted mean-squared error
where \(g_x\) is the spectral density of the regressor process. Cagan’s regression is precisely such a constrained fit: with \(y_t=m_t-p_t\), the fitted class is the one-parameter geometric-lag family \(b^1(L;\alpha)=\alpha\,\dfrac{1-\lambda}{1-\lambda L}\), and the true projection (495) is the member of that family with coefficient \(\operatorname{plim}\hat\alpha\). Least squares therefore drives \(\hat\alpha\) to the value minimizing (497), which is exactly (496). The “approximation error” \(\operatorname{plim}\hat\alpha-\alpha\) is the gap between the structural coefficient and the spectral-density-weighted best fit — non-zero because, under the equilibrium money rule, the regressor is correlated with the error.
Reading the bias through (497) makes two things vivid. First, the weighting density \(g_x\) is the spectrum of inflation, which is generated by the money-supply rule (490): change the monetary regime and \(g_x\) changes, so the “structural” coefficient Cagan estimates is not invariant across regimes. Sims’s approximation formula here wears the clothes of Lucas’s critique — the same regime-dependence of a fitted decision/demand rule that runs through Some Applications to Rational Expectations Models, Chapter XIV, and Exact Linear Rational Expectations Models. Second, it locates the inconsistency where it belongs: not in the shape \(\dfrac{1-\lambda}{1-\lambda L}\) (which Cagan gets right) but in the interpretation of a projection coefficient as a structural elasticity.
Consistent estimation, identification, and findings#
Because inflation and money growth are determined simultaneously, recovering \(\alpha\) requires a system method. The bivariate model (491) can be written as a vector ARMA(1,1) whose innovations \(a_t=(a_{1t},a_{2t})'\) are the one-step-ahead forecast errors of \((x_t,\mu_t)\); crucially, \(\alpha\) does not enter the innovation recursions, so a Gaussian full-information maximum-likelihood estimator (Wilson 1973) identifies \(\lambda\) and the innovation covariance \(D_a=(\sigma_{11},\sigma_{12},\sigma_{22})\) by minimizing \(\lvert\hat D_a(\lambda)\rvert\) over the single parameter \(\lambda\). But the mapping from the four identified quantities \((\lambda,\sigma_{11},\sigma_{12},\sigma_{22})\) to the five structural parameters \((\alpha,\lambda,\sigma_\varepsilon^2,\sigma_\eta^2,\sigma_{\varepsilon\eta})\) is not invertible: \(\lambda\) and \(\sigma_\varepsilon^2\) are identified, but \(\alpha\) and \(\sigma_{\varepsilon\eta}\) are not separately identified — offsetting changes in the two leave the likelihood unchanged. To estimate \(\alpha\) at all one must impose a restriction; Sargent sets \(\sigma_{\varepsilon\eta}=0\) (money-supply and portfolio shocks uncorrelated), which yields an estimator of \(\alpha\) that depends sensitively on the estimated covariance matrix of the forecast errors and must be regarded as delicate. (The QuantEcon lecture implements the estimator and, when \(\sigma_{\varepsilon\eta}=0\), an equivalent instrumental-variables procedure that uses the fitted inflation innovations as an instrument for expected inflation.)
The quantitative findings, summarized from the paper’s tables, bear out the “statistical artifact” reading of the paradox:
The estimates of \(\alpha\) are very loose. For most hyperinflations the maximum-likelihood standard error on \(\hat\alpha\) is of the same order as the point estimate; only Hungary I is estimated with much precision. For Hungary I, \(\hat\alpha\approx-1.84\) implies a revenue-maximizing inflation of \(-1/\hat\alpha\approx54\%\) per month against an observed \(46\%\) — the one case where the estimate substantially weakens Cagan’s paradox. For the others the point estimates do not eliminate the paradox, but two-standard-error bands comfortably include values of \(\alpha\) that would.
The one-parameter model survives overfitting for several countries. Testing the restricted representation — whose systematic part contains the single free parameter \(\lambda\) — against six vector-ARMA overparameterizations by likelihood-ratio \(\chi^2\) statistics, the model is not rejected at the 5% level for Germany, Greece, and Poland; Hungary I and Austria are rejected by several parameterizations, Russia by one. That a representation with one systematic free parameter survives at all is striking.
The demand for money during hyperinflations may not have been as sharply isolated as Cagan’s tight estimates suggested, and the slope of the portfolio-balance schedule is difficult or impossible to estimate precisely under the money-supply regimes that actually prevailed. The apparent paradox is largely what Sims’s approximation-error formula predicts a misspecified least-squares regression will produce.
References#
Theodore W. Anderson. The Statistical Analysis of Time Series. John Wiley & Sons, New York, 1971.
Robert J. Barro. Inflation, the payments period, and the demand for money. Journal of Political Economy, 78(6):1228–1263, 1970.
Phillip Cagan. The monetary dynamics of hyperinflation. In Milton Friedman, editor, Studies in the Quantity Theory of Money, pages 25–117. University of Chicago Press, Chicago, 1956.
Clive W. J. Granger. Investigating causal relations by econometric models and cross-spectral methods. Econometrica, 37(3):424–438, 1969.
John F. Muth. Optimal properties of exponentially weighted forecasts. Journal of the American Statistical Association, 55(290):299–306, 1960.
Thomas J. Sargent. The demand for money during hyperinflations under rational expectations: i. International Economic Review, 18(1):59–82, 1977.
Thomas J. Sargent and Neil Wallace. Rational expectations and the dynamics of hyperinflation. International Economic Review, 14(2):328–350, 1973.
Christopher A. Sims. Money, income, and causality. The American Economic Review, 62(4):540–552, 1972.
Christopher A. Sims. The role of approximate prior restrictions in distributed lag estimation. Journal of the American Statistical Association, 67(337):169–175, 1972.
G. Tunnicliffe Wilson. The estimation of parameters in multivariate time series models. Journal of the Royal Statistical Society, Series B, 35(1):76–85, 1973.