Estimating Dynamic Equilibrium Models with Stochastic Volatility by 六 Jesús Fernández-Villaverde ** Pablo Guerrón-Quintana *** Juan F. Rubio-Ramírez
Documento de Trabajo 2013-23
December 2013
University of Pennsylvania, NBER, CEPR, and FEDEA. ** Federal Reserve Bank of Philadelphia *** Duke University, Federal Reserve Bank of Atlanta, and FEDEA.
Los Documentos de Trabajo se distribuyen gratuitamente a las Universidades e Instituciones de Investigación que lo solicitan. No obstante están disponibles en texto completo a través de Internet: http://www.fedea.es. These Working Paper are distributed free of charge to University Department and other Research Centres. They are also available through Internet: http://www.fedea.es. ISSN:1696-750
Estimating Dynamic Equilibrium Models with Stochastic Volatility
Jes˙s Fern·ndez-Villaverdey Pablo GuerrÛn-Quintanaz
Juan F. Rubio-RamÌrezx
May 8, 2013
yUniversity of Pennsylvania, NBER, CEPR, and FEDEA, <jesusfv@econ.upenn.edu>.
We thank AndrÈ Kurmann, Jim Nason, Frank Schorfheide, Neil Shephard, Tao Zha, and participants at several seminars for useful comments, and BÈla SzemÈly for invaluable research assistance. This paper borrows material from our manuscript ìFortune or Virtue: Time-Variant Volatilities Versus Parameter Drifting in U.S. Data.î Beyond the usual disclaimer, we must note that any views expressed herein are those of the authors and not necessarily those of the Federal Reserve Bank of Atlanta, the Federal Reserve Bank of Philadelphia, or the Federal Reserve System. Finally, we also thank the NSF for Önancial support. This paper is available free of charge at http://www.philadelphiafed.org/researchand-data/publications/working-papers/.
zFederal Reserve Bank of Philadelphia, <pablo.guerron@phil.frb.org>.
xDuke University, Federal Reserve Bank of Atlanta, and FEDEA, <juan.rubio-ramirez@duke.edu>.
Abstract
We propose a novel method to estimate dynamic equilibrium models with stochastic volatility. First, we characterize the properties of the solution to this class of models. Second, we take advantage of the results about the structure of the solution to build a sequential Monte Carlo algorithm to evaluate the likelihood function of the model. The approach, which exploits the profusion of shocks in stochastic volatility models, is versatile and computationally tractable even in large-scale models, such as those often employed by policy-making institutions. As an application, we use our algorithm and Bayesian methods to estimate a business cycle model of the U.S. economy with both stochastic volatility and parameter drifting in monetary policy. Our application shows the importance of stochastic volatility in accounting for the dynamics of the data.
Keywords: Dynamic equilibrium models, Stochastic volatility, Parameter drifting, Bayesian methods.
JEL classiÖcation numbers: E10, E30, C11.
1. Introduction
Over the last few years, economists have uncovered the importance of time-varying volatility in accounting for the dynamics of U.S. aggregate data. Stock and Watson (2002) and Sims and Zha (2006) are some prominent examples of this line of research. As a response, a growing literature has developed to investigate the origins and consequences of time-varying volatility. A particularly successful branch of this literature uses dynamic equilibrium models to emphasize the role of volatility shocks (also known as uncertainty shocks, but in any case speciÖed as stochastic volatility as in Shephard, 2008) in business cycle áuctuations. Among those we can highlight are Bloom et al. (2007), Fern·ndez-Villaverde and Rubio-RamÌrez (2007), Justiniano and Primiceri (2008), Bloom (2009), Fern·ndez-Villaverde et al. (2010c), and Bloom et al. (2012). In these models, there are two types of shocks: structural shocks (shock to productivity, to preferences, to policy, etc.) and volatility shocks (shocks to the standard deviation of the innovations to the structural shocks).
To fulÖll the promise in this literature, it is crucial to have a set of tools to take these dynamic equilibrium models with stochastic volatility to the data. Namely, we want to estimate, evaluate, and compare these models using a likelihood-based approach. However, the task is complicated by the inherent non-linearity that stochastic volatility generates. By construction, linearization is ill-equipped to handle time-varying volatility models because it yields certainty-equivalent policy functions. That is, volatility ináuences neither the agentsídecision rules nor the laws of motion of the aggregate variables. Hence, to consider how volatility a§ects those, it becomes imperative to employ a non-linear solution method for the dynamics of the economy. Once we consider non-linear solution methods we are forced to use simulation-based estimators of the likelihood.
In principle, and beyond the computational burden, this could seem a mild challenge. For instance, one could be tempted to get around this problem by solving the model non-linearly with a fast algorithm and employing, for estimation, a particle Ölter to get an unbiased simulationbased estimator of the likelihood. The particle Ölter presented in Fern·ndez-Villaverde and Rubio-RamÌrez (2007) would be a natural candidate for such an undertaking. Unfortunately, that version of the particle Ölter relies, when estimating models with stochastic volatility, on the presence of linear measurement errors in observables. Otherwise, we would be forced to solve a large quadratic system of equations with multiple solutions, an endeavor for which there are no suitable algorithms. Although measurement errors are both plausible and empirically relevant, their presence complicates the interpretation of the results. Thus, it is desirable to consider a particle Ölter that departs from requiring measurement errors.
To accomplish this goal, we show how to write a particle Ölter that allows us to evaluate the likelihood of a dynamic equilibrium model with stochastic volatility by exploiting the structure of the second-order approximation to the decision rules that characterize the equilibrium of the economy without the need of measurement errors. Second-order approximations capture the implications of stochastic volatility and are convenient because they are accurate but not computationally expensive.
We achieve the objective in two steps. First, we characterize the second-order approximation to the decision rules of a rather general dynamic equilibrium model with stochastic volatility.1 Second, we demonstrate how we can use this characterization to handily simulate the likelihood function of the model with a particle Ölter without measurement errors. The key is to show how the decision rules characterization reduces the quadratic problem associated with the evaluation of the approximated measurement density to a much simpler matrix inversion problem. This is a novel extension of sequential Monte Carlo methods that can be useful for other applications. After we have evaluated the likelihood, we can combine it with a prior and an Markov chain Monte Carlo (McMc) algorithm to draw from the posterior distribution (Flury and Shephard, 2011).
Our characterization of the second-order approximation to the decision rules should be of interest in itself even for researchers who are not evaluating the likelihood of their model but who are keen to understand how the model works. Among others, the results are useful to analyze the equilibrium of the model, to explore the shape of its impulse-response functions, or to calibrate it.
More concretely, we prove that:
1. The Örst-order approximation to the decision rules of the agents (or any other equilibrium object of interest) does not depend on volatility shocks and they are certainty equivalent.
2. The second-order approximation to the decision rules of the agents only depends on volatility shocks on terms where volatility is multiplied by the innovation to its own structural shock. For instance, if we have two structural shocks (a productivity and a preference shock) and a volatility shock to each of them, the only non-zero term where the volatility shock to productivity appears is the one where the volatility shock to productivity multiplies the innovation to the productivity shock. Thus, of the terms in the second-order approximation that complicate the evaluation of the likelihood function, only a few are non-zero.
1These theorems generalize previous results derived by Schmitt-GrohÈ and Uribe (2004), for the homoscedastic shocks case to the time-varying case.
3. The perturbation parameter will only appear in a non-zero term where it is raised to a square. This term is a constant that corrects for precautionary behavior induced by risk.
As an application, we estimate a dynamic equilibrium model of the U.S. economy with stochastic volatility using the particle Ölter and Bayesian methods. The model, an otherwise standard business cycle model with nominal rigidities, incorporates not only stochastic volatility in the shocks that drive its dynamics but also parameter drifting in the parameters that control monetary policy. In that way, we include two of the main mechanisms that researchers have highlighted to account for the time-varying volatility of U.S. data: heteroscedastic shocks and parameter drifting. We include these together in the model because a ìone-at-a-timeî investigation is fraught with peril. If we allow only one source of variation in the model, the likelihood may want to take advantage of this extra degree of áexibility to Öt the data better. For example, if the ìtrueîmodel is one with parameter drifting in monetary policy, an estimated model without drift but with stochastic volatility in one of the structural shocks may interpret the ìtrueî drift as time-varying volatility in that shock. If, instead, we had time-varying volatility in technological shocks in the data, an estimated model without stochastic volatility and only parameter drifting may conclude, erroneously, that the parameters of monetary policy are changing. Lastly, we have a model that is as rich as the models employed by policy-making institutions in real life. While solving and estimating such a large model is a computational challenge, we wanted to demonstrate that our procedure is of practical use and to make our application a blueprint for the estimation of other dynamic equilibrium models.
Our main empirical Öndings are as follows. First, the posterior distribution of parameters puts most of its mass in areas that denote a fair amount of stochastic volatility, yet another proof of the importance of time-varying volatility in applied models. Second, a model comparison exercise indicates that, even after controlling for stochastic volatility, the data clearly prefer a speciÖcation where monetary policy changes over time. This Önding should not be interpreted, though, as implying that volatility shocks did not play a role of their own. It means, instead, that a successful model of the U.S. economy requires the presence of both stochastic volatility and parameter drifting, a result that challenges the results of Sims and Zha (2006). Finally, we document the evolution of the structural shocks, of stochastic volatility, and the parameters of monetary policy. We emphasize the conáuence, during the 1970s, of times of high volatility and weak responses to ináation, and, during the 1990s, of positive structural shocks and low volatility even if monetary policy was weaker than often argued. In the appendix, we thoroughly explore the model, including the construction of counterfactual histories of the U.S. data by varying some aspect of the model such as shutting down time-varying volatility or imposing alternative monetary policies.
An alternative to our stochastic volatility framework would be to work with Markov regimeswitching models such as those of Bianchi (2009) or Farmer et al. (2009). This class of models provides an extra degree of áexibility in modelling aggregate dynamics that is highly promising. In fact, some of the fast changes in policy parameters that we document in our empirical section suggest that discrete jumps could be a good representation of the data. However, current technical limitations regarding the computation of the equilibria induced by regime switches force researchers to focus on small models that are only stylized representations of an economy.
Finally, even if the motivation for our approach and the application belong to macroeconomics, there is nothing speciÖc about that Öeld in the tools we present. One can think about the importance of estimating dynamic equilibrium models with stochastic volatility in many other Öelds, from Önance (Bansal and Yaron, 2004, would be a prominent example) to industrial organization (for instance, an equilibrium model of industry dynamics where the demand and supply shocks have time-varying volatility), international trade (where innovations to the real exchange rates or to the country spreads are well known to have time-varying volatility, see Fern·ndez-Villaverde et al., 2010c), or many others.
The rest of the paper is organized as follows. Section 2 presents a generic dynamic equilibrium model with stochastic volatility to Öx notation and discuss how to solve it. Section 3 explains the evaluation of the likelihood of the model. Section 4 introduces a business cycle model as an application of our procedure, takes it to the U.S. data, and comments on the main Öndings. Section 5 concludes. An extensive technical appendix includes details about the proofs of the main results in the paper, the model, the computation, and additional empirical exercises.
2. Dynamic Equilibrium Models with Stochastic Volatility
2.1. The Model
The set of equilibrium conditions of a wide class of dynamic equilibrium models can be written as
\[\mathbb {E} _ {t} f \left(\mathcal {Y} _ {t + 1}, \mathcal {Y} _ {t}, \mathcal {S} _ {t + 1}, \mathcal {S} _ {t}, \mathcal {Z} _ {t + 1}, \mathcal {Z} _ {t}; \gamma\right) = 0,\tag{1}\]
where is the conditional expectation operator at time t, is the vector of observables at time t, is the vector of endogenous states at time is the vector of structural shocks at time maps into , and is the vector of parameters that describe preferences and technology. In the context of this paper, is also the vector of parameters to be estimated.
We will consider models where the structural shocks, , follow a stochastic volatility process of the form
\[\mathcal {Z} _ {i t + 1} = \rho_ {i} \mathcal {Z} _ {i t} + \Lambda \sigma_ {i} \sigma_ {i t + 1} \varepsilon_ {i t + 1}\tag{2}\]
for all , where is a perturbation parameter, is the mean volatility, and log the percentage deviation of the standard deviation of the innovations to the structural shocks with respect to its mean, evolves as
\[\log \sigma_ {i t + 1} = \vartheta_ {i} \log \sigma_ {i t} + \Lambda \left(1 - \vartheta_ {i} ^ {2}\right) ^ {\frac {1}{2}} \eta_ {i} u _ {i t + 1}\tag{3}\]
for all . The combination of levels in (2) and logs in (3) ensures a positive We multiply the innovation in (3) by to normalize its size by the persistence of . It will be clear momentarily why we specify (2) and (3) in terms of the perturbation parameter . It is also convenient to write, for all , the laws of motions for and log
\[\mathcal {Z} _ {i t} = \rho_ {i} \mathcal {Z} _ {i t - 1} + \sigma_ {i} \sigma_ {i t} \varepsilon_ {i t}\tag{4}\]
and
\[\log \sigma_ {i t} = \vartheta_ {i} \log \sigma_ {i t - 1} + (1 - \vartheta_ {i} ^ {2}) ^ {\frac {1}{2}} \eta_ {i} u _ {i t}\tag{5}\]
Note that the perturbation parameter appears only in equations (2) and (3) but not in equations (4) and (5). This is because this parameter is used to eliminate, when we later determine the point around which to perform a higher-order approximation to the equilibrium dynamics of the model, any uncertainty about the future. Given the information set in equation (1), there is uncertainty about both and log for all , but there is no uncertainty about either or log for any
Let be the vector of volatility shocks, the vector of innovations to the structural shocks, and the vector of innovations to the volatility shocks. We assume that and , where 0 is an vector of zeros and I is an identity matrix. To ease notation, we assume that:
1. All structural shocks face volatility shocks and that the volatility shocks are uncorrelated. It is straightforward, yet cumbersome, to generalize the notation to other cases.
2. and are normally distributed. As we will see below, to implement our particle Ölter, we only need to be able to simulate and evaluate the density of .
2.2. The Solution
Given equations (1)-(5), the solution to the model -if one exists- that embodies equilibrium dynamics (that is, agentís optimization and market clearing conditions) is characterized by a policy function describing the evolution of the endogenous state variables
\[\mathcal {S} _ {t + 1} = h \left(\mathcal {S} _ {t}, \mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}, \mathcal {E} _ {t}, \mathcal {U} _ {t}, \Lambda ; \gamma\right),\tag{6}\]
and two policy functions describing the law of motion of the observables
\[\mathcal {Y} _ {t} = g \left(\mathcal {S} _ {t}, \mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}, \mathcal {E} _ {t}, \mathcal {U} _ {t}, \Lambda ; \gamma\right)\tag{7}\]
and
\[\mathcal {Y} _ {t + 1} = g \left(\mathcal {S} _ {t + 1}, \mathcal {Z} _ {t}, \Sigma_ {t}, \Lambda \mathcal {E} _ {t + 1}, \Lambda \mathcal {U} _ {t + 1}, \Lambda ; \gamma\right),\tag{8}\]
together with equations (4) and (5) describing the laws of motion of the structural and volatility shocks. The policy functions h and g map into and ; respectively, and are indexed by the vector of parameters .
For our purposes, it is important to deÖne the steady state of the model. Our assumptions about the stochastic processes imply that, in the steady state, and . Given a vector of parameters and equations (1)-(8), the steady state of the model is a vector of observables and an vector of endogenous states such that
\[f \left(g \left(h \left(\mathcal {S}, \mathbf {0}, \mathbf {0}, \mathbf {0}, \mathbf {0}, 0; \gamma\right), \mathbf {0}, \mathbf {0}, \mathbf {0}, \mathbf {0}, 0; \gamma\right), g \left(\mathcal {S}, \mathbf {0}, \mathbf {0}, \mathbf {0}, \mathbf {0}, 0; \gamma\right), h \left(\mathcal {S}, \mathbf {0}, \mathbf {0}, \mathbf {0}, \mathbf {0}, 0; \gamma\right), \mathcal {S}, \mathbf {0}, \mathbf {0}; \gamma\right) = 0.\tag{9}\]
In addition, in the steady state, the following two relationships hold
\[\mathcal {S} = h \left(\mathcal {S}, \mathbf {0}, \mathbf {0}, \mathbf {0}, \mathbf {0}, 0; \gamma\right),\tag{10}\]
and
\[\mathcal {Y} = g \left(\mathcal {S}, \mathbf {0}, \mathbf {0}, \mathbf {0}, \mathbf {0}, 0; \gamma\right).\tag{11}\]
The perturbation parameter plays a key role in deÖning our steady state. If ; the model is in the steady state, since we eliminate any uncertainty about the future. If ; the model is not in the steady state and conditions (9)-(11) do not hold. For instance, in general because of the precautionary behavior of agents.
2.3. The State-Space Representation
Given the solution to the model -the policy functions (6)-(8) together with equations (4) and (5)- we can concisely characterize the equilibrium dynamics of the model by its state-space representation written in terms of the transition and the observation equations.
The transition equation uses the policy function (6), together with equations (4) and (5), to describe the evolution of the states (endogenous states, structural shocks, volatility shocks, and their innovations) as a function of lag states, the perturbation parameter, and the vector of parameters
\[\mathbb {S} _ {t + 1} = \widetilde {h} \left(\mathbb {S} _ {t}, \Lambda ; \gamma\right) + \Xi \mathbb {W} _ {t + 1}\tag{12}\]
where is the vector of the states and maps into . Also, is a vector of random variables, and are vectors with distributions and , and is a matrix with the the top rows equal to zero and the bottom matrix equal to the identity matrix. Note that and share distributions with and respectively. If we were to change distributions for either or we would need to do the same with and . Let us deÖne . The linearity of the second term of the right-hand side is a consequence of the autoregressive speciÖcation of the structural shocks and the evolution of their volatilities. However, through the function we let those shocks and their volatilities a§ect nonlinearly.
The measurement equation uses the policy function (7) to describe the relationship of the observables with the states, the perturbation parameter, and the vector of parameters
\[\mathcal {Y} _ {t} = g \left(\mathbb {S} _ {t}, \Lambda ; \gamma\right).\tag{13}\]
2.4. Approximating the State-Space Representation
In general, when we deal with the class of dynamic equilibrium models with stochastic volatility, we do not have access to a closed-form solution, that is, the policy functions h and g cannot be found explicitly. Thus, we cannot build the state-space representation described by (12) and (13). Instead, we will approximate numerically the solution of model and use the result to generate an approximated state-space representation.
There are many di§erent solution algorithms for dynamic equilibrium models. Among those, the perturbation method has emerged as a popular way (see Judd and Guu, 1997, Schmitt-GrohÈ and Uribe, 2004, and Aruoba et al., 2006 among others) to obtain higher-order Taylor series approximations to the policy functions (6)-(7) together with equations (4) and (5) around the steady state. It is key to note that we also get the Taylor expansion of (4) because it is a nonlinear function. Beyond being extremely fast for systems with a large number of state variables, perturbation o§ers high levels of accuracy even relatively far away from the perturbation point (Aruoba, et al., 2006, and, for cases with stochastic volatility, Caldara et al., 2012).
The perturbation method allows the researcher to approximate the solution to the model up to any order. Since we want to analyze models with volatility shocks, we must go beyond a Örst-order approximation. First-order approximations are certainty equivalent and hence, they are silent about stochastic volatility. As we well see in section 3.1, volatility shocks only appear starting in the second-order approximation. But since we want to evaluate the likelihood function, we need to stop at a second-order approximation. Because of dimensionality issues, higher-order approximation would make the evaluation of the likelihood function exceedingly challenging for current computers for models with a reasonable number of state variables.
Given a second-order approximation to the policy functions (6)-(7) together with equations (4) and (5) around the steady state, the approximated state-space representation can be written in terms of two equations: the approximated transition equation and the approximated measurement equation. The approximated transition equation is
\[\widehat {\mathbb {S}} _ {t + 1} = \left( \begin{array}{c} \Psi_ {s, 1} ^ {1} \widehat {\mathbb {S}} _ {t} \\ \vdots \\ \Psi_ {s, n _ {s}} ^ {1} \widehat {\mathbb {S}} _ {t} \end{array} \right) + \frac {1}{2} \left( \begin{array}{c} \widehat {\mathbb {S}} _ {t} ^ {\prime} \Psi_ {s, 1} ^ {2} \widehat {\mathbb {S}} _ {t} \\ \vdots \\ \widehat {\mathbb {S}} _ {t} ^ {\prime} \Psi_ {s, n _ {s}} ^ {2} \widehat {\mathbb {S}} _ {t} \end{array} \right) + \frac {1}{2} \left( \begin{array}{c} \Psi_ {s, 1} ^ {\Lambda} \\ \vdots \\ \Psi_ {s, n _ {s}} ^ {\Lambda} \end{array} \right) + \Xi \mathbb {W} _ {t + 1}\tag{14}\]
where is the vector of deviations of the states from their steady-state value and . The approximated measurement equation is
\[\mathcal {Y} _ {t} - \mathcal {Y} = \left( \begin{array}{c} \Psi_ {y, 1} ^ {1} \widehat {\mathbb {S}} _ {t} \\ \vdots \\ \Psi_ {y, k} ^ {1} \widehat {\mathbb {S}} _ {t} \end{array} \right) + \frac {1}{2} \left( \begin{array}{c} \widehat {\mathbb {S}} _ {t} ^ {\prime} \Psi_ {y, 1} ^ {2} \widehat {\mathbb {S}} _ {t} \\ \vdots \\ \widehat {\mathbb {S}} _ {t} ^ {\prime} \Psi_ {y, k} ^ {2} \widehat {\mathbb {S}} _ {t} \end{array} \right) + \frac {1}{2} \left( \begin{array}{c} \Psi_ {y, 1} ^ {\Lambda} \\ \vdots \\ \Psi_ {y, k} ^ {\Lambda} \end{array} \right).\tag{15}\]
where is the steady-state value of .
In these equations, is a vector and is an matrix for The Örst term is the linear approximation to transition equation for the states, while the second term is the quadratic component of the second-order approximation. Similarly, is a vector and an matrix for . The interpretation of each term is the same as before: the Örst term is the linear component and the second one the quadratic component of the approximation to the law of motion for the measurement equations. The term is the constant that appears in the second-order perturbation and that is often interpreted as the correction for risk in the evolution of state . Similarly, the term is the constant correction for risk of observable . All these vectors and matrices are non-linear functions of the vector of parameters . It is important to emphasize that we are not assuming the presence of any measurement error. We will return to this point in a few paragraphs.
3. Stochastic Volatility and Evaluation of the Likelihood
In this section, we explain how to evaluate the likelihood function of our class of dynamic equilibrium models with volatility shocks in equation (1) using the approximated transition and the measurement equations (14) and (15). If we allow to be the data counterpart of our observables (with to be their history up to time t, given a vector of parameters , we can write the likelihood of as
\[\prod_ {t = 1} ^ {T} p \left(\mathcal {Y} _ {t} = \mathbb {Y} _ {t} | \mathbb {Y} ^ {t - 1}; \gamma\right),\]
where
\[\begin{array}{c} p \left(\mathcal {Y} _ {t} = \mathbb {Y} _ {t} | \mathbb {Y} ^ {t - 1}; \gamma\right) = \\ \int \int \int \int p \left(\mathcal {Y} _ {t} = \mathbb {Y} _ {t} | \mathcal {S} _ {t}, \mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}, \mathcal {E} _ {t}; \gamma\right) p \left(\mathcal {S} _ {t}, \mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}, \mathcal {E} _ {t} | \mathbb {Y} ^ {t - 1}; \gamma\right) d \mathcal {S} _ {t} d \mathcal {Z} _ {t - 1} d \Sigma_ {t - 1} d \mathcal {E} _ {t} \end{array}\tag{16}\]
for all and
\[\begin{array}{c} p \left(\mathcal {Y} _ {1} = \mathbb {Y} _ {1}; \gamma\right) \\ = \int \int \int \int p \left(\mathcal {Y} _ {1} = \mathbb {Y} _ {1} | \mathcal {S} _ {1}, \mathcal {Z} _ {0}, \Sigma_ {0}, \mathcal {E} _ {1}; \gamma\right) p \left(\mathcal {S} _ {1}, \mathcal {Z} _ {0}, \Sigma_ {0}, \mathcal {E} _ {1}; \gamma\right) d \mathcal {S} _ {1} d \mathcal {Z} _ {0} d \Sigma_ {0} d \mathcal {E} _ {1}. \end{array}\tag{17}\]
Note that the ís do not show up in these two expressions. It will be momentarily clear why this comes directly from our procedure below to evaluate the likelihood.
Computing this likelihood is a di¢ cult problem. Since we do not have analytic forms for the terms inside the integral, it cannot be evaluated exactly and deterministic numerical integration algorithms are too slow for practical use (we have four integrals per period over large dimensions). As a feasible alternative, we will show how to use a simple particle Ölter to obtain a simulationbased estimate of (16) for the class of dynamic equilibrium models with volatility shocks we described above. K¸nsch (2005) proves, under weak conditions, that the particle Ölter delivers a consistent estimator of the likelihood function and that a central limit theorem applies. A particle Ölter is a sequential simulation device for Öltering of nonlinear and/or non-Gaussian space models (Pitt and Shephard, 1999, and Doucet et al., 2001). In economics the particle Ölter has been used, among others, by Fern·ndez-Villaverde and Rubio-RamÌrez (2007).
As mentioned before, the particle Ölter has minimal requirements: the ability to evaluate the approximated measurement density associated with the approximated measurement equation, to simulate from the approximated dynamics of the state using the approximated transition equation, and to draw from the unconditional density of the states implied by the approximated transition equation. Usually, the Örst requirement is the hardest. These three requirements are formally described in the following assumption.
Assumption 1. To implement the particle Ölter, we assume that:
1. We can evaluate the approximated measurement density
\[p \left(\mathcal {Y} _ {t} = \mathbb {Y} _ {t} | \mathcal {S} _ {t}, \mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}, \mathcal {E} _ {t}; \gamma\right).\]
2. We can simulate from the approximated transition equation
\[\left(\mathcal {S} _ {t + 1} ^ {\prime}, \mathcal {Z} _ {t} ^ {\prime}, \Sigma_ {t} ^ {\prime}, \mathcal {E} _ {t + 1} ^ {\prime}\right) ^ {\prime} | \left(\mathcal {S} _ {t} ^ {\prime}, \mathcal {Z} _ {t - 1} ^ {\prime}, \Sigma_ {t - 1} ^ {\prime}, \mathcal {E} _ {t} ^ {\prime}\right) ^ {\prime}, \mathcal {F} \left(\mathbb {Y} ^ {t}\right); \gamma\]
for all , where is the Öltration of .
3. We can draw from the unconditional distribution implied by the approximated transition equation
\[p \left(\mathcal {S} _ {t}, \mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}, \mathcal {E} _ {t}; \gamma\right).\]
The second requirement asks for the Öltration of , because we will need to evaluate the volatility shocks (this will be clearer below). The last requirement can be easily implemented using the results in Santos and Peralta-Alva (2005). Given our assumption about and the quadratic form of the approximated transition equation (14), the second requirement easily holds. A key novelty of this paper is that we show how the class of dynamic equilibrium models with volatility shocks considered here also satisÖes the Örst requirement.
For our class of models, conditional on having N draws of from the density , each integral (16) can be consistently approximated by
\[p \left(\mathcal {Y} _ {t} = \mathbb {Y} _ {t} | \mathbb {Y} ^ {t - 1}; \gamma\right) \simeq \frac {1}{N} \sum_ {i = 1} ^ {N} p \left(\mathcal {Y} _ {t} = \mathbb {Y} _ {t} | s _ {t} ^ {i}, z _ {t - 1} ^ {i}, \sigma_ {t - 1} ^ {i}, \varepsilon_ {t} ^ {i}; \gamma\right)\tag{18}\]
for all . For , we need N draws from the density , so that the integral (17) can be consistently approximated by
\[p \left(\mathcal {Y} _ {1} = \mathbb {Y} _ {1}; \gamma\right) \simeq \frac {1}{N} \sum_ {i = 1} ^ {N} p \left(\mathcal {Y} _ {1} = \mathbb {Y} _ {1} | s _ {1} ^ {i}, z _ {0} ^ {i}, \sigma_ {0} ^ {i}, \varepsilon_ {1} ^ {i}; \gamma\right).\tag{19}\]
We know that we can make the draws because requirement in assumption 1 holds.
In our framework, checking whether requirement 1 in assumption 1 holds means checking whether we can evaluate the approximated measurement density
\[p \left(\mathcal {Y} _ {t} = \mathbb {Y} _ {t} | s _ {t} ^ {i}, z _ {t - 1} ^ {i}, \sigma_ {t - 1} ^ {i}, \varepsilon_ {t} ^ {i}; \gamma\right)\tag{20}\]
for each draw in for all . This evaluation step is crucial, not only because it is required to compute (18) and (19) but also because, as we will explain momentarily, our particle Ölter needs to evaluate (20) to resample from and get draws from in order to obtain for all in a recursive way.
To check how we can evaluate the approximated measurement density (20), we rewrite (15) in terms of draws and , instead of and . Thus, the approximated measurement equation (15) can be rewritten as
\[\mathbb {Y} _ {t} - \mathcal {Y} = \left( \begin{array}{c} \Psi_ {y, 1} ^ {1} \widehat {\mathbb {S}} _ {t} ^ {i} \\ \vdots \\ \Psi_ {y, k} ^ {1} \widehat {\mathbb {S}} _ {t} ^ {i} \end{array} \right) + \frac {1}{2} \left( \begin{array}{c} \widehat {\mathbb {S}} _ {t} ^ {i \prime} \Psi_ {y, 1} ^ {2} \widehat {\mathbb {S}} _ {t} ^ {i} \\ \vdots \\ \widehat {\mathbb {S}} _ {t} ^ {i \prime} \Psi_ {y, k} ^ {2} \widehat {\mathbb {S}} _ {t} ^ {i} \end{array} \right) + \frac {1}{2} \left( \begin{array}{c} \Psi_ {y, 1} ^ {\Lambda} \\ \vdots \\ \Psi_ {y, k} ^ {\Lambda} \end{array} \right)\tag{21}\]
where and . The new approximated measurement equation (21) implies that evaluating the approximated measurement density (20) involves solving the system of quadratic equations
\[\mathbb {Y} _ {t} - \mathcal {Y} - \left( \begin{array}{c} \Psi_ {y, 1} ^ {1} \widehat {\mathbb {S}} _ {t} ^ {i} \\ \vdots \\ \Psi_ {y, k} ^ {1} \widehat {\mathbb {S}} _ {t} ^ {i} \end{array} \right) - \frac {1}{2} \left( \begin{array}{c} \widehat {\mathbb {S}} _ {t} ^ {i \prime} \Psi_ {y, 1} ^ {2} \widehat {\mathbb {S}} _ {t} ^ {i} \\ \vdots \\ \widehat {\mathbb {S}} _ {t} ^ {i \prime} \Psi_ {y, k} ^ {2} \widehat {\mathbb {S}} _ {t} ^ {i} \end{array} \right) - \frac {1}{2} \left( \begin{array}{c} \Psi_ {y, 1} ^ {\Lambda} \\ \vdots \\ \Psi_ {y, k} ^ {\Lambda} \end{array} \right) = 0\tag{22}\]
for given ; and
Solving this system is non-trivial. Since the system is quadratic, we may have either none or several di§erent solutions. In fact, it is even hard to know how many solutions there are. But even if we knew the number of solutions, we are not aware of any accurate and e¢ cient method to solve quadratic problems that Önds all the solutions. This di¢ culty seemingly prevents us from achieving our goal of evaluating the likelihood function.
A solution would be to introduce, as Fern·ndez-Villaverde and Rubio-RamÌrez (2007) did, a k 1 vector of linear measurement errors and solve for those instead of . In this case the system would have a unique, easy to Önd solution. However, there are three reasons, in increasing order of importance, why this solution is not satisfactory:
1. Although measurement errors are both plausible and empirically relevant, their presence complicates the interpretation of any empirical results. In particular, we are interested in measuring how heteroscedastic structural shocks help in accounting for the data and measurement errors can confound us.
2. The absence of measurement errors will help us to illustrate below how dynamic equilibrium models with volatility shocks have a profusion of shocks that we can exploit in our estimation.
3. As we will also show below, volatility shocks would enter linearly (conditional on the draw) in the system equations. Since, by deÖnition, linear measurement errors would also enter in the same fashion, it would be hard to identify one apart from the others.
Our alternative in this paper is to realize that considering stochastic volatility converts the above-described quadratic system into a linear one. Hence, if a rank condition is satisÖed, the system (22) has only one solution and the solution can be found by simply inverting a matrix. Thus, requirement 1 in assumption 1 holds and we can use the particle Ölter. The core of the argument is to note that, when volatility shocks are considered, the policy functions share a peculiar pattern that we can exploit.
3.1. Structure of the Solution
Our Örst step will characterize the Örst- and second-order derivatives of the policy functions h and g evaluated at the steady state. Then, we will describe an interesting pattern in these derivatives. The second step will take advantage of the pattern to show that, when the number of volatility shocks equals the number of observables, our quadratic system becomes a linear one. This characterization is important both for estimation and, more generally, for the analysis of perturbation solutions to dynamic equilibrium models with stochastic volatility.
3.1.1. First- and Second-order Derivatives of the Policy Functions
Let us begin with the characterization of the Örst- and second-order derivatives of the policy functions. The following theorem shows that the Örst-order derivatives of h and g with respect to any component of and evaluated at the steady state are zero; that is, volatility shocks and their innovations do not a§ect the linear component of the optimal decision rule of the agents. The same occurs with the perturbation parameter . A similar result has been established by Schmitt-GrohÈ and Uribe (2004) for the homoscedastic shocks case.
Theorem 2. Let us denote as the derivative of the element of generic function with respect to the element of generic variable ! evaluated at the non-stochastic steady state (where we drop this index if ! is unidimensional). Then, for the dynamic equilibrium model speciÖed in equation (1), we have that for , and
Proof. See appendix 6.1.1.
The second theorem shows, among other things, that the second partial derivatives of h and g with respect to either log and any other variable but is also zero for any
Theorem 3. Furthermore, if we denote as the derivative of the element of generic function with respect to the element of generic variable and the element of generic variable evaluated at the non-stochastic steady state (where again we drop the index for unidimensional variables), we have that , and
\[\left[ h _ {\Lambda , \mathcal {Z} _ {t - 1}} \right] _ {j} ^ {i _ {1}} = \left[ g _ {\Lambda , \mathcal {Z} _ {t - 1}} \right] _ {j} ^ {i _ {2}} = \left[ h _ {\Lambda , \Sigma_ {t - 1}} \right] _ {j} ^ {i _ {1}} = \left[ g _ {\Lambda , \Sigma_ {t - 1}} \right] _ {j} ^ {i _ {2}} = [ h _ {\Lambda , \mathcal {E} _ {t}} ] _ {j} ^ {i _ {1}} = [ g _ {\Lambda , \mathcal {E} _ {t}} ] _ {j} ^ {i _ {2}} = [ h _ {\Lambda , \mathcal {U} _ {t}} ] _ {j} ^ {i _ {1}} = [ g _ {\Lambda , \mathcal {U} _ {t}} ] _ {j} ^ {i _ {2}} = 0\]
for , and ,
\[\left[ h _ {\mathcal {S} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\mathcal {S} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = [ h _ {\mathcal {S} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = [ g _ {\mathcal {S} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = 0\]
for
\[\left[ h _ {\mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = \left[ h _ {\Sigma_ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\Sigma_ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = 0\]
and
\[\left[ h _ {\mathcal {Z} _ {t - 1}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\mathcal {Z} _ {t - 1}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = \left[ h _ {\Sigma_ {t - 1}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\Sigma_ {t - 1}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = [ h _ {\mathcal {U} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = [ g _ {\mathcal {U} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = 0\]
for , and , and
\[\left[ h _ {\mathcal {E} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\mathcal {E} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = \left[ h _ {\mathcal {E} _ {t}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\mathcal {E} _ {t}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = 0\]
for i1 2 f1; : : : ; ng, i2 2 f1; : : : ; kg, and
Proof. See appendix 6.1.2.
Since the statement of theorem 3 is long and involved, we clarify it with a table in which we characterize the second derivatives of the functions h and g with respect to the di§erent variables . This pattern is both interesting and useful for our purposes.
Table 3.1: Second Derivatives
| $\mathcal{S}_{t}\mathcal{S}_{t} \neq 0$ | $\mathcal{S}_{t}\mathcal{Z}_{t-1} \neq 0$ | $\mathcal{S}_{t}\Sigma_{t-1} = 0$ | $\mathcal{S}_{t}\mathcal{E}_{t} \neq 0$ | $\mathcal{S}_{t}\mathcal{U}_{t} = 0$ | $\mathcal{S}_{t}\Lambda = 0$ |
| $\mathcal{Z}_{t-1}\mathcal{Z}_{t-1} \neq 0$ | $\mathcal{Z}_{t-1}\Sigma_{t-1} = 0$ | $\mathcal{Z}_{t-1}\mathcal{E}_{t} \neq 0$ | $\mathcal{Z}_{t-1}\mathcal{U}_{t} = 0$ | $\mathcal{Z}_{t-1}\Lambda = 0$ | |
| $\Sigma_{t-1}\Sigma_{t-1} = 0$ | $\Sigma_{t-1}\mathcal{E}_{t} \neq 0^{*}$ | $\Sigma_{t-1}\mathcal{U}_{t} = 0$ | $\Sigma_{t-1}\Lambda = 0$ | ||
| $\mathcal{E}_{t}\mathcal{E}_{t} \neq 0$ | $\mathcal{E}_{t}\mathcal{U}_{t} \neq 0^{*}$ | $\mathcal{E}_{t}\Lambda = 0$ | |||
| $\mathcal{U}_{t}\mathcal{U}_{t} = 0$ | $\mathcal{U}_{t}\Lambda = 0$ | ||||
| $\Lambda\Lambda \neq 0$ |
The way to read table 3.1 is as follows. Take an arbitrary entry, for instance, entry (1,2), . In this entry, we state that the cross-derivatives of h and with respect to and are di§erent from zero (the table is upper triangular because, given the symmetry of second derivatives, we do not need to report those entries). Similarly, entry (3,5), tells us that the cross-derivatives of h and g with respect to and are all zero.2 Entries (3,4) and (4,5)
2 For ease of exposition, in table 3.1 we are not being explicit about the dimensions of the matrices: it is a purely qualititive description of the relevant derivatives.
have a ì*î to denote that the only cross-derivatives of those entries that are di§erent from zero are those that correspond to the same index j (that is, the cross-derivatives of each innovation to the structural shocks with respect to its own volatility shock and the cross-derivatives of the innovation to the structural shocks to the innovation to its own volatility shock). The lower triangular part of the table is empty because of the symmetry of second derivatives.
Table 3.1 tells us that, of the 21 possible sets of second derivatives, 12 are zero and 9 are not. The implications for the decision rules of agents and for the equilibrium function are striking. The perturbation parameter, , will only have a coe¢ cient di§erent from zero in the term where it appears in a square by itself. This term is a constant that corrects for precautionary behavior induced by risk. Volatility shocks, , appear with coe¢ cients di§erent from zero only in the term where they are multiplied by the innovation to its own structural shock . Finally, innovations to the volatility shocks, , also appear with coe¢ cients di§erent from zero when they show up with the innovation to their own structural shock . Hence, of the terms that complicate the evaluation of the approximated measurement density, only the ones associated with and are non-zero.
3.1.2. Evaluating the Likelihood Using the Particle Filter
The second step is to use theorems 2 and 3 to show that the system (22) is linear and it has only one solution. Corollary 4 shows that the pattern described in table 3.1. has an important, yet rather direct implication for the structure of the approximated measurement equation (21).
Corollary 4. The approximated measurement equation (21) can be written as
\[\mathbb {Y} _ {t} - \mathcal {Y} = \left( \begin{array}{c} \widetilde {\Psi} _ {y, 1} ^ {1} \widetilde {\mathbb {S}} _ {t} ^ {i} \\ \vdots \\ \widetilde {\Psi} _ {y, k} ^ {1} \widetilde {\mathbb {S}} _ {t} ^ {i} \end{array} \right) + \frac {1}{2} \left( \begin{array}{c} \widetilde {\widehat {\mathbb {S}}} _ {t} ^ {i \prime} \widetilde {\Psi} _ {y, 1} ^ {2, 1} \widetilde {\mathbb {S}} _ {t} ^ {i} \\ \vdots \\ \widetilde {\widehat {\mathbb {S}}} _ {t} ^ {i \prime} \widetilde {\Psi} _ {y, k} ^ {2, 1} \widetilde {\mathbb {S}} _ {t} ^ {i} \end{array} \right) + \frac {1}{2} \left( \begin{array}{c} \Psi_ {y, 1} ^ {\Lambda} \\ \vdots \\ \Psi_ {y, k} ^ {\Lambda} \end{array} \right) + \left( \begin{array}{c} \varepsilon_ {t} ^ {i \prime} \widetilde {\Psi} _ {y, 1} ^ {2, 2} \\ \vdots \\ \varepsilon_ {t} ^ {i \prime} \widetilde {\Psi} _ {y, k} ^ {2, 2} \end{array} \right) \sigma_ {t - 1} ^ {i} + \left( \begin{array}{c} \varepsilon_ {t} ^ {i \prime} \widetilde {\Psi} _ {y, 1} ^ {2, 3} \\ \vdots \\ \varepsilon_ {t} ^ {i \prime} \widetilde {\Psi} _ {y, k} ^ {2, 3} \end{array} \right) \mathcal {U} _ {t}\]
where is the vector of the simulated states without the stochastic volatility components (that is, endogenous states, structural shocks, and their innovations), is its steady state, and is its deviation from the steady state. Let us deÖne . The matrix is a vector that represents the Örst-order component of the second-order approximation to the law of motion for the measurement equation as a function of , where and , for . The matrix is an matrix that represents the second-order component of the second-order approximation to the law of motion for the measurement equation as a function of for . The matrix is an matrix that represents the second-order component of the second-order approximation to the law of motion for the measurement equation as a function of and for The matrix is an matrix that represents the second-order term of the approximated law of motion for the measurement equation as a function of and for
We are now ready to show that the system (22) is linear, and if a rank condition is satisÖed, it has only one solution. DeÖne
\[\begin{array}{c} \mathbb {A} \left(\mathbb {Y} _ {t}, s _ {t} ^ {i}, z _ {t - 1} ^ {i}, \sigma_ {t - 1} ^ {i}, \varepsilon_ {t} ^ {i}; \gamma\right) \\ \equiv \mathbb {Y} _ {t} - \mathcal {Y} - \left( \begin{array}{c} \widetilde {\Psi} _ {y, 1} ^ {1} \widetilde {\widehat {\mathbb {S}}} _ {t} ^ {i} \\ \vdots \\ \widetilde {\Psi} _ {y, k} ^ {1} \widetilde {\widehat {\mathbb {S}}} _ {t} ^ {i} \end{array} \right) - \frac {1}{2} \left( \begin{array}{c} \widetilde {\widehat {\mathbb {S}}} _ {t} ^ {i \prime} \widetilde {\Psi} _ {y, 1} ^ {2, 1} \widetilde {\widehat {\mathbb {S}}} _ {t} ^ {i} \\ \vdots \\ \widetilde {\widehat {\mathbb {S}}} _ {t} ^ {i \prime} \widetilde {\Psi} _ {y, k} ^ {2, 1} \widetilde {\widehat {\mathbb {S}}} _ {t} ^ {i} \end{array} \right) - \frac {1}{2} \left( \begin{array}{c} \Psi_ {y, 1} ^ {\Lambda} \\ \vdots \\ \Psi_ {y, k} ^ {\Lambda} \end{array} \right) - \left( \begin{array}{c} \varepsilon_ {t} ^ {i \prime} \widetilde {\Psi} _ {y, 1} ^ {2, 2} \\ \vdots \\ \varepsilon_ {t} ^ {i \prime} \widetilde {\Psi} _ {y, k} ^ {2, 2} \end{array} \right) \sigma_ {t - 1} ^ {i} \end{array}\]
and
\[\mathbb {B} \left(\varepsilon_ {t} ^ {i}; \gamma\right) \equiv \left( \begin{array}{c c c} \left(\varepsilon_ {t} ^ {i \prime} \widetilde {\Psi} _ {y, 1} ^ {2, 3}\right) ^ {\prime} & \ldots & \left(\varepsilon_ {t} ^ {i \prime} \widetilde {\Psi} _ {y, k} ^ {2, 3}\right) ^ {\prime} \end{array} \right) ^ {\prime}.\]
Let , then is a matrix. If B is full rank, the solution to the system (22) can be written as
\[\mathcal {U} _ {t} = \mathbb {B} ^ {- 1} \left(\varepsilon_ {t} ^ {i}; \gamma\right) \mathbb {A} \left(\mathbb {Y} _ {t}, s _ {t} ^ {i}, z _ {t - 1} ^ {i}, \sigma_ {t - 1} ^ {i}, \varepsilon_ {t} ^ {i}; \gamma\right).\]
The next theorem shows how to use this solution to evaluate the approximated measurement density.
Theorem 5. Let , then B is a matrix. is full rank, the approximated measurement density can be written as
\[p \left(\mathcal {Y} _ {t} = \mathbb {Y} _ {t} | s _ {t} ^ {i}, z _ {t - 1} ^ {i}, \sigma_ {t - 1} ^ {i}, \varepsilon_ {t} ^ {i}; \gamma\right) = \det \mathbb {B} ^ {- 1} \left(\varepsilon_ {t} ^ {i}; \gamma\right) p \left(\mathcal {U} _ {t} = \mathbb {B} ^ {- 1} \left(\varepsilon_ {t} ^ {i}; \gamma\right) \mathbb {A} \left(\mathbb {Y} _ {t}, s _ {t} ^ {i}, z _ {t - 1} ^ {i}, \sigma_ {t - 1} ^ {i}, \varepsilon_ {t} ^ {i}; \gamma\right)\right)\tag{23}\]
for each draw in for all , which can be evaluated given that we know , and the distribution of
Proof. The theorem is a straightforward application of the change of variables theorem.
Theorem 5 shows that requirement 1 in assumption 1 holds. Consequently, we can use a variation of the particle Ölter adapted for our class of dynamic equilibrium models with volatility shocks. We are requiring to be full rank. When can not be full rank? would fail to have full rank when the impact of volatility innovations are identical across several elements of . This would mean that the observables lack enough information to tell volatility shocks apart. If this is the case, a new set of observables would have to be chosen to estimate the model.
Note also that theorem 5 assumes . Given the notation in section 2.1, this means that the number of structural shocks equals to the number of observables. This is not always necessary. What the theorem needs is that the number of volatility shocks equals the number of observables. Since, to simplify notation, we have assumed that all structural shocks face volatility shocks, the number of structural shocks equals the number of observables. As mentioned in section 2.1, we could have structural shocks that do not face volatility shocks (this will be the case in our application to follow). In that case, we could have more structural shocks than observables. From the theorem, it is also clear that if we used rather than to compute the approximated measurement equation, we would have to solve a quadratic system, a challenging task.
An outline of the algorithm in pseudo-code is:
Step 0: Set . Sample N values from Step 1: Compute
\[p \left(\mathcal {Y} _ {t} = \mathbb {Y} _ {t} | \mathbb {Y} ^ {t - 1}; \gamma\right) \simeq \frac {1}{N} \sum_ {i = 1} ^ {N} \det \mathbb {B} ^ {- 1} \left(\varepsilon_ {t} ^ {i}; \gamma\right) p \left(\mathcal {U} _ {t} = \mathbb {B} ^ {- 1} \left(\varepsilon_ {t} ^ {i}; \gamma\right) \mathbb {A} \left(\mathbb {Y} _ {t}, s _ {t} ^ {i}, z _ {t - 1} ^ {i}, \sigma_ {t - 1} ^ {i}, \varepsilon_ {t} ^ {i}; \gamma\right)\right)\]
using expression (23) and the importance weights for each draw
\[q _ {t} ^ {i} = \frac {\det \mathbb {B} ^ {- 1} \left(\varepsilon_ {t} ^ {i} ; \gamma\right) p \left(\mathcal {U} _ {t} = \mathbb {B} ^ {- 1} \left(\varepsilon_ {t} ^ {i} ; \gamma\right) \mathbb {A} \left(\mathbb {Y} _ {t} , s _ {t} ^ {i} , z _ {t - 1} ^ {i} , \sigma_ {t - 1} ^ {i} , \varepsilon_ {t} ^ {i} ; \gamma\right)\right)}{\sum_ {i = 1} ^ {N} \det \mathbb {B} ^ {- 1} \left(\varepsilon_ {t} ^ {i} ; \gamma\right) p \left(\mathcal {U} _ {t} = \mathbb {B} ^ {- 1} \left(\varepsilon_ {t} ^ {i} ; \gamma\right) \mathbb {A} \left(\mathbb {Y} _ {t} , s _ { t } ^ {i} , z _ { t - 1 } ^ {i} , \sigma_ {t - 1} ^ {i} , \varepsilon_ {t} ^ {i} ; \gamma\right)\right)}.\]
Step 2: Sample N times with replacement from and probability . This delivers
Step 3: Simulate from the approximated transition equation
\[\left(\mathcal {S} _ {t + 1} ^ {\prime}, \mathcal {Z} _ {t} ^ {\prime}, \Sigma_ {t} ^ {\prime}, \mathcal {E} _ {t + 1} ^ {\prime}\right) ^ {\prime} | \left(s _ {t | t} ^ {i \prime}, z _ {t - 1 | t} ^ {i \prime}, \sigma_ {t - 1 | t} ^ {i \prime}, \varepsilon_ {t | t} ^ {i \prime}\right) ^ {\prime}, \mathcal {F} \left(\mathbb {Y} ^ {t}\right); \gamma .\]
Step 4: If ; set and go to step 1. Otherwise stop.
Once we have evaluated the likelihood, we can nest this algorithm either with a McMc to perform Bayesian inference (as done in our application; see Flury and Shephard, 2011, for technical details) or with some optimization algorithm to undertake maximum likelihood estimation (as done, in a model without volatility shocks, by Van Binsbergen et al., 2012). In this last case, care must be taken to keep the random numbers used in the particle Ölter constant among iterations to achieve equi-continuity and to use an optimization algorithm that does not rely on derivatives, as the particle Ölter implies an evaluation of the likelihood function that is not di§erentiable.
4. An Application: A Business Cycle Model with Stochastic Volatility
As an illustration of our procedure, we estimate a business cycle model of the U.S. economy with nominal rigidities and stochastic volatility. We will show: 1) how we can characterize posterior distributions of the parameters of interest and 2) how we can recover and analyze the smoothed structural and volatility shocks. Those are two of the most relevant exercises in terms of the estimation of dynamic equilibrium models. Once the estimation has been undertaken, there are many other exercises that can be done with the empirical results. For instance, in the appendix, we include several of those such as: 1) Önding the impulse response functions (IRFs) of the model; 2) evaluating the Öt of the model in comparison with some alternatives; and 3) undertaking some counterfactuals and run alternative histories of the evolution of the U.S. economy.
Before doing though, we need to motivate our choice of application. The model that we present has many strengths but also important weaknesses. Su¢ ce it to say here that since this model is the base of much applied policy analysis by academics and policy-making institutions, it is a natural example for this paper. Thus, given that the model is well known, our presentation will be brief. We will depart only along two dimensions from the standard model: we will have stochastic volatility in the shocks that drive the dynamics of the economy (the object of interest in this paper) and parameter drifting in the monetary policy rule. In that way, the likelihood has the chance of picking between two of the alternatives that the literature has highlighted to account for the well-documented time-varying volatility of the U.S. aggregate time series, either a reduced volatility of the shocks that hit the economy (as in Sims and Zha, 2006) or a di§erent monetary policy (Clarida et al., 2000), which makes the application of interest in itself.
4.1. Households
The economy is populated by a continuum of households indexed by and preferences:
\[\mathbb {E} _ {0} \sum_ {t = 0} ^ {\infty} \beta^ {t} d _ {t} \left\{\log \left(c _ {j t} - h c _ {j t - 1}\right) + v \log \left(\frac {m _ {j t}}{p _ {t}}\right) - \varphi_ {t} \psi \frac {l _ {j t} ^ {1 + \vartheta}}{1 + \vartheta} \right\},\]
which are separable in consumption, , real money balances, ; and hours worked, . In our notation, is the conditional expectations operator, is the discount factor, h controls habit persistence, is the inverse of the Frisch labor supply elasticity, is a shifter to intertemporal preference that follows log where ; and is a labor supply shifter that evolves as log where : The two preference shocks are common to all households.
The principal novelty of these preferences is that, for both shifters and the standard deviations, and , of their innovations, and , are indexed by time; that is, they stochastically move according to:
\[\log \sigma_ {d t} = \rho_ {\sigma_ {d}} \log \sigma_ {d t - 1} + \left(1 - \rho_ {\sigma_ {d}} ^ {2}\right) ^ {\frac {1}{2}} \eta_ {d} u _ {d t} \mathrm{where} u _ {d t} \sim \mathcal {N} (0, 1)\]
and
\[\log \sigma_ {\varphi t} = \rho_ {\sigma_ {\varphi}} \log \sigma_ {\varphi t - 1} + \left(1 - \rho_ {\sigma_ {\varphi}} ^ {2}\right) ^ {\frac {1}{2}} \eta_ {\varphi} u _ {\varphi t} \text {where} u _ {\varphi t} \sim \mathcal {N} (0, 1).\]
This parsimonious speciÖcation introduces only four new parameters, , and , while being surprisingly powerful in capturing important features of the data (Shephard, 2008). All the shocks and innovations throughout the model are observed by the agents when they are realized. Agents have, as well, rational expectations about how they evolve over time.
We can interpret the shocks to preferences and to their volatility as reáecting the random evolution of more complicated phenomena. For example, stochastic volatility may appear as the consequence of changing demographics. An economy with older agents might be both less patient because of higher mortality risk (in our notation, a lower and less prone to reallocations in the labor force because of longer attachments to particular jobs (in our notation, a lower
Although we assume complete Önancial markets, to ease notation we drop the Arrow securities implied by that assumption from the budget constraints (they are in net zero supply at the aggregate level anyway). Households also hold government bonds that pay a nominal gross interest rate of . Therefore, the householdís budget constraint is given by:
\[c _ {j t} + x _ {j t} + \frac {m _ {j t}}{p _ {t}} + \frac {b _ {j t + 1}}{p _ {t}} = w _ {j t} l _ {j t} + \left(r _ {t} u _ {j t} - \mu_ {t} ^ {- 1} \Phi [ u _ {j t} ]\right) k _ {j t - 1} + \frac {m _ {j t - 1}}{p _ {t}} + R _ {t - 1} \frac {b _ {j t}}{p _ {t}} + T _ {t} + F _ {t}\]
where is investment, is the real wage, is the real rental price of capital, is the rate of use of capital, is the cost of using capital at rate in terms of the Önal good, is an investment-speciÖc technological level, is a lump-sum transfer, and is ÖrmsíproÖts. We specify that , a form that satisÖes that 2 and . This function carries the normalization that in the balanced growth path of the economy. Using the relevant Örst-order conditions, we can Önd where is the (rescaled) steady-state rental price of capital (determined by the other parameters in the model). This leaves us with only one free parameter,
The capital accumulated by household j at the end of period t is given by:
\[k _ {j t} = (1 - \delta) k _ {j t - 1} + \mu_ {t} \left(1 - \frac {\kappa}{2} \left(\frac {x _ {j t}}{x _ {j t - 1}} - \Lambda_ {x}\right) ^ {2}\right) x _ {j t}\]
where is the depreciation rate and is an investment adjustment cost parameter. This function is written in deviations with respect to the balanced growth rate of investment, . The investmentspeciÖc technology level , follows a random walk in logs, log with and where is the drift of the process and is the innovation to its growth rate. The standard deviation also evolves as:
\[\log \sigma_ {\mu t} = \rho_ {\sigma_ {\mu}} \log \sigma_ {\mu t - 1} + \left(1 - \rho_ {\sigma_ {\mu}} ^ {2}\right) ^ {\frac {1}{2}} \eta_ {\mu} u _ {\mu t} \text {where} u _ {\mu t} \sim \mathcal {N} (0, 1).\]
Again, we can interpret this stochastic volatility as a stand-in for a more detailed explanation of technological progress in capital production that we do not model explicitly.
Each household j supplies a di§erent type of labor services that are aggregated by a labor packer into homogeneous labor with the production function that is rented to intermediate good producers at wage . The labor packer is perfectly competitive and it takes all wages as given. Households set their wages with a Calvo pricing mechanism. At the start of every period, a randomly selected fraction of households can reoptimize their wages (where, by a law of large numbers, individual probabilities and aggregate fractions are equal). All other households index their wages given past ináation with an indexation parameter
4.2. Firms
There is one Önal good producer that aggregates a continuum of intermediate goods according to the production function:
\[y _ {t} = \left(\int_ {0} ^ {1} y _ {i t} ^ {\frac {\varepsilon - 1}{\varepsilon}} d i\right) ^ {\frac {\varepsilon}{\varepsilon - 1}}\tag{24}\]
where is the elasticity of substitution. The Önal good producer is perfectly competitive and minimizes its costs subject to the production function (24) and taking as given all intermediate goods prices and the Önal good price
Each intermediate good is produced by a monopolistic competitor with technology ; where is the capital rented by the Örm, is the amount of the ìpackedî labor input rented by the Örm, and is neutral productivity. Productivity evolves as log where is the drift of the process and is the innovation to its growth rate. The time-varying standard deviation of this innovation follows:
\[\log \sigma_ {A t} = \rho_ {\sigma_ {A}} \log \sigma_ {A t - 1} + \left(1 - \rho_ {\sigma_ {A}} ^ {2}\right) ^ {\frac {1}{2}} \eta_ {A} u _ {A t} \text {where} u _ {A t} \sim \mathcal {N} (0, 1).\]
Intermediate good producers meet the quantity demanded by the Önal good producer by renting and at prices and . Given their demand function, these producers set prices to maximize proÖts. However, when they do so, they follow a Calvo pricing scheme. In each period, a fraction of intermediate good producers reoptimize their prices. All other Örms partially index their prices by past ináation with an indexation parameter
4.3. The Monetary Authority
The model is closed with a monetary authority that sets the nominal interest rates by following a modiÖed Taylor rule:
\[\frac {R _ {t}}{R} = \left(\frac {R _ {t - 1}}{R}\right) ^ {\gamma_ {R}} \left(\left(\frac {\Pi_ {t}}{\Pi}\right) ^ {\gamma_ {\Pi} \gamma_ {\Pi , t}} \left(\frac {y _ {t} ^ {d}}{y _ {t - 1} ^ {d}} / \exp (\Lambda_ {y ^ {d}})\right) ^ {\gamma_ {y} \gamma_ {y, t}}\right) ^ {1 - \gamma_ {R}} \xi_ {t}.\tag{25}\]
The Örst term on the right-hand side represents a desire for interest rate smoothing, expressed in terms of the balanced growth path nominal interest rate. The second term, an ìináation gap,î responds to the deviation of ináation from its balanced growth path level . The third term is a ìgrowth gapî: the ratio between the growth rate of the economy and , the balanced path gross growth rate of , where is aggregate demand (deÖned precisely in appendix 6.2). The last term is the monetary policy shock, where log with an innovation and a time-varying standard deviation, , that follows an autoregressive process
\[\log \sigma_ {\xi t} = \rho_ {\sigma_ {\xi}} \log \sigma_ {\xi t - 1} + \left(1 - \rho_ {\sigma_ {\xi}} ^ {2}\right) ^ {\frac {1}{2}} \eta_ {\xi} u _ {\xi , t} \text {where} u _ {\xi , t} \sim \mathcal {N} (0, 1).\]
In this policy rule, we have two drifting parameters: the responses of the monetary authority to the ináation gap and the growth gap. The parameters drift over time as:
\[\log \gamma_ {\Pi t} = \rho_ {\gamma_ {\Pi}} \log \gamma_ {\Pi t - 1} + \sigma_ {\pi} \varepsilon_ {\pi t} \text {and} \log \gamma_ {y t} = \rho_ {\gamma_ {y}} \log \gamma_ {y t - 1} + \sigma_ {y} \varepsilon_ {y t} \text {where} \varepsilon_ {\pi t}, \varepsilon_ {y t} \sim \mathcal {N} (0, 1).\]
For simplicity, the volatility of the innovation to this processes is Öxed over time. The agents perfectly observe the changes in monetary policy parameters.
4.4. Equilibrium and Solution
We can characterize the equilibrium of the model in appendix 6.2. This equilibrium conditions are non-stationary because we have two unit roots in the processes for technology. However, we circumvent this problem by rescaling the model using the variable in the form: , and . In the notation of section 2.1 we have that:
1. The states of the (rescaled) economy are
\[\mathcal {S} _ {t} = \left(\log \widetilde {k} _ {t - 1}, \log \widetilde {c} _ {t - 1}, \log \widetilde {x} _ {t - 1}, \log \widetilde {y} _ {t - 1}, \log v _ {t - 1} ^ {p}, \log v _ {t - 1} ^ {w}, \log \widetilde {w} _ {t - 1}, \log R _ {t - 1}, \log \Pi_ {t - 1}\right) ^ {\prime}.\]
2. The structural shocks are . The parameter drifts are handled as structural shocks in the state-space representation.
3. The volatility shocks are ; where the last two zeros correspond to the processes for parameter drifting, which have constant volatilities (see also in vector below). Here, we can see that we do not need as many volatility shocks (5 of them) as structural shocks (7 of them).
4. The innovations to the structural and the volatility shocks are and ; respectively.
We pick as observables the Örst di§erence of the log of the relative price of investment, the log federal funds rate, log ináation, the Örst di§erence of log output, and the Örst di§erence of log real wages, in our notation : We select these variables because they bring us information about aggregate behavior (output), the stand of monetary policy (the interest rate and ináation), and the di§erent shocks (the relative price of investment about investment-speciÖc technological change, the other four variables about technology and preference shocks) that we are concerned about. Note that we have the same number of observables as we do of volatility shocks, as required by theorem 5.
We follow the steps in section 2.2 and employ a second-order perturbation around the rescaled steady-state to approximate this equilibrium and build the associated state-space representation (appendix 6.4 gives more details). The second-order approximation is even more relevant in our application because, to complicate matters, a linearization would also imply that the parameter drift in the Taylor rule would disappear as well from the equilibrium dynamics (see appendix 6.3).
4.5. Data and Estimation
We estimate our model using the Öve time series for the U.S. economy described above. Our sample covers 1959.Q1 to 2007.Q1, with 192 observations. Because of space considerations, we stop at 2007 to avoid having to deal with the recent Önancial crisis, which would make it di¢ cult to appreciate the points we want to illustrate about how to econometrically deal with stochastic volatility. This could be Öxed at the cost of a lengthier discussion. Appendix 6.5 explains how we construct the series.
Once we have evaluated the likelihood, we combine it with a prior. We pick áat priors on a bounded support for all the parameters. The bounds are either natural economic restrictions (for instance, the Calvo and indexation parameters lie between 0 and 1) or are so wide that the likelihood assigns (numerically) zero probability to values outside them. Bounded áat priors induce a proper posterior, a convenient feature for our exercises below.
We resort to áat priors for two reasons. First, to reduce the impact of presample information and show that our results arise mainly from the shape of the likelihood and not from the prior (although, of course, áat priors are not invariant to reparameterization). Thus, the reader who wants to interpret our posterior modes as maximum likelihood point estimates can do so. Second, because as we learned in Fern·ndez-Villaverde et al. (2010c), eliciting priors for stochastic volatility is di¢ cult, since we deal with unfamiliar units, such as the variance of volatility shocks, about which we do not have clear beliefs. Flat priors come, though, at a price: before proceeding to the estimation, we have to Öx several parameters to reduce the dimensionality of the problem.
Table 4.1: Fixed Parameters
| $\beta$ | h | $\psi$ | $\vartheta$ | $\delta$ | $\alpha$ | $\kappa$ | $\varepsilon$ | $\eta$ | $\Phi_{2}$ | $\rho_{\gamma_{\Pi}}$ | $\rho_{\gamma_{y}}$ | $\sigma_{y}$ |
| 0.99 | 0.9 | 8 | 1.17 | 0.025 | 0.21 | 9.5 | 10 | 10 | 0.001 | 0.95 | 0 | 0 |
Table 4.1 lists the Öxed parameters. Our guiding criterion in selecting them was to pick conventional values in the literature. The discount factor, , is a default choice, habit persistence, ; matches the observed sluggish response of consumption to shocks, the parameter controlling the level of labor supply, ; captures the average amount of hours in the data, and the depreciation rate, , induces the appropriate capital-output ratio. The elasticities of substitution, , deliver average mark-ups of around 10 percent, a common value in these models. We set the cost of capital utilization, , to a small number to introduce some curvature in this decision. Three parameter values are borrowed from the point estimates from a similar model without stochastic volatility or parameter drifting presented in Fern·ndez-Villaverde et al. (2009). The Örst is the inverse of the Frisch labor elasticity, . As argued by Chetty et al. (2011), this aggregate elasticity is compatible with micro data, once we allow for intensive and extensive margins on labor supply. The second is the coe¢ cient of the intermediate goods production function, . This value is lower than the common calibration of Cobb-Douglas production functions in real business cycle models because, in our environment, we have positive proÖts that appear as capital income in the National Income and Product Accounts. Finally, the adjustment cost, , is in line with other estimates from similar models ( would be particularly hard to identify, since investment is not one of our observables).
The autoregressive parameter of the evolution of the response to ináation, , is set to 0.95. In preliminary estimations, we discovered that the likelihood pushed this parameter to 1. When this happened, the simulations became numerically unstable: after a series of positive innovations to log , the reaction of nominal interest rates to ináation could be too tepid for too long. The 0.95 value seems to be the highest possible value of such that the problem does not appear. The last two parameters, and are equal to zero because, also in exploratory estimations, the likelihood favored values of : Thus, we decided to forget about them and make
To Önd the posterior, we proceed as follows. First, we deÖne a grid of parameter values and check for the regions of high posterior density by evaluating the likelihood function in each point of the grid. This is a time-consuming procedure, but it ensures that we are searching in the right zone of the parameter space. Once we have identiÖed the global maximum in the grid, we initialize a random-walk Metropolis-Hastings algorithm from this point. After an extensive Öne-tuning of the algorithm, we use 10,000 draws from the chain to compute posterior moments.
4.6. Results I: Parameter Estimates
Our Örst empirical result is the parameter estimates. To ease the discussion, we group them in di§erent tables, one for each set of parameters dealing with related aspects of the model. In all cases, we report the mode of the posterior and the standard deviation in parenthesis below (in the interest of space, we do not include the whole histograms of the posterior).
Table 4.2: Posterior, Parameters of Nominal Rigidities and Structural Shocks
| $\theta_p$ | $\chi$ | $\theta_w$ | $\chi_w$ | $\rho_d$ | $\rho_\varphi$ | $\Lambda_\mu$ | $\Lambda_A$ |
| 0.8139(0.0143) | 0.6186(0.024) | 0.6869(0.0432) | 0.6340(0.0074) | 0.1182(0.0049) | 0.9331(0.0425) | 0.0034(6.6e-5) | 0.0028(4.1e-5) |
Table 4.2 presents the results for the nominal rigidities and the stochastic processes for the structural shocks parameters. Our estimates indicate an economy with substantial rigidities in prices, which are reoptimized roughly once every Öve quarters, and in wages, which are reoptimized approximately every three quarters. Moreover, since the standard deviations are small, there is enough information in the data about this result. At the same time, there is a fair amount of indexation, between 0.62-0.63, which brings a strong persistence of ináation. While it is tempting to compare our estimates with the micro evidence on the individual duration of prices, in our model all prices and wages change every quarter. That is why, to a naive observer, our economy would look like one displaying tremendous price áexibility.
We estimate a low persistence of the intertemporal preference shock and a high persistence of the intratemporal one. The low estimate of produces the quick variations in marginal utilities of consumption that match output growth and ináation áuctuations. The intratemporal shock is persistent to account for long-lived movements in hours worked. We estimate mean growth rates of technology of 0.0034 (neutral) and 0.0028 (investment-speciÖc). Those numbers give us an average growth of the economy of 0.44 percent per quarter, or around 1.77 percent on an annual basis (0.46 and 1.86 percent in the data, respectively). Technology shocks, in our model, are deviations with respect to these drifts. Thus, we estimate that falls in only 8 of the 192 quarters in our sample (which roughly corresponds to the percentage of quarters where measured productivity falls in the data), even if we estimate negative innovations to neutral technology in 103 quarters.
Table 4.3: Posterior, Parameters of the Stochastic Processes for Volatility Shocks
| $\log \sigma_d$ | $\log \sigma_\varphi$ | $\log \sigma_\mu$ | $\log \sigma_A$ | $\log \sigma_\xi$ |
| -1.9834(0.0726) | -2.4983(0.0917) | -6.0283(0.1278) | -3.9013(0.0745) | -6.000(0.1471) |
| $\rho_{\sigma_d}$ | $\rho_{\sigma_\varphi}$ | $\rho_{\sigma_\mu}$ | $\rho_{\sigma_a}$ | $\rho_{\sigma_\xi}$ |
| 0.9506(0.0298) | 0.1275(0.0032) | 0.7508(0.035) | 0.2411(0.005) | 0.8550(0.0231) |
| $\eta_d$ | $\eta_\varphi$ | $\eta_\mu$ | $\eta_a$ | $\eta_\xi$ |
| 0.1007(0.0083) | 2.8316(0.0669) | 0.3115(0.006) | 0.7720(0.013) | 0.5723(0.0185) |
The results for the parameters of the stochastic volatility processes appear in table 4.3. In all cases, the and the are far away from zero: the likelihood strongly favors values where stochastic volatility plays an important role. The standard deviations of the innovations of the intertemporal preference shock and of the monetary policy shock are the most persistent, while the standard deviation of the innovation of the intratemporal preference shock is the least persistent. The standard deviation of the innovations of the volatility shock to the intratemporal preference shock, , is large: the model asks for fast changes in the size of movements in marginal utilities of leisure to reproduce the hours data.
Table 4.4: Posterior, Policy Parameters
| $\gamma_R$ | $\log \gamma_y$ | $\Pi$ | $\log \gamma_{\Pi}$ | $\sigma_\pi$ |
| 0.7855(0.0162) | -1.4034(0.0498) | 1.0005(0.0043) | 0.0441(0.0005) | 0.145(0.002) |
In table 4.4., we have the estimates of the policy parameters. The autoregressive component of the federal funds rate is high, 0:7855, although somewhat smaller than in estimations without parameter drift. The value of (0.24 in levels) is similar to other results in the literature and shows that the likelihood clearly likes parameter drifting, although with mild persistence. The estimated value of plus the correction on equilibrium ináation implied by second-order e§ects of the solution match the average ináation in the data.3 Finally, the estimated value of (1.045 in levels) guarantees local determinacy of the equilibrium even if is temporarily below 1 (see appendix 6.6 for details).
In appendix 6.7 we plot the impulse response functions of the model implied by our estimates. This exercise allows us to check that the estimates are sensible and that the behavior of the model is consistent with the behavior of related models in the literature. In appendix 6.8 we compare our model against an alternative version without parameter drifting but still with stochastic volatility. That is, we ask whether, once we have included stochastic volatility, it is still important to allow for changes in the monetary policy rule to account for the time-varying volatility of U.S. aggregate data over the last several decades. The results show that, even after controlling for stochastic volatility, the data strongly prefer a speciÖcation where the monetary policy rule has changed over time. In appendix 6.9 we show, though, that this Önding does not imply that volatility shocks did not play an important role in the time-varying volatilities of U.S. aggregate time series. In other words, both stochastic volatility and parameter drifting are key parts of a successful dynamic equilibrium model of the U.S. economy.
3Also, these second-order e§ects enormously complicate the introduction of time-variation in . The likelihood wants to match the moments of the ergodic distribution of ináation, not the level of , which is ináation along the balanced growth path. When we have non-linearities, the mean of that ergodic distribution may be far from : Thus, learning about is hard. Learning about a time-varying is even harder.
4.7. Results II: Smoothed Shocks
Figure 4.1 reports the log-deviations with respect to their means for the smoothed intertemporal, intratemporal, and monetary shocks and deviations of the growth rate of the investment and technological shocks with respect to their means ( 2 standard deviations). To ease reading of the results, we color di§erent vertical bars to represent each of the periods at the Federal Reserve: the McChesney Martin years from the start of our sample in 1959 to the appointment of Burns in February 1970 (white), the Burns-Miller era (light blue), the Volcker interlude from August 1979 to August 1987 (grey), the Greenspan times (orange), and Bernankeís tenure from February 2006 (yellow).





Figure 4.1: Smoothed intertemporal shock, intratemporal shock, investment-speciÖc shock, technology shock, and monetary policy shock +/- 2 s.d.

We see in the top left panel of Ögure 4.1 that the intertemporal shock, log is particularly high in the 1970s. This increases householdsídesire for current consumption (for instance, because of the entrance of baby boomers into adulthood). A higher aggregate demand triggers, in the model, the higher ináation observed in the data for those years. The shock has a dramatic drop in the second quarter of 1980. This is precisely the quarter in which the Carter administration invoked the Credit Control Act (March 14, 1980). Schreft (1990) documents that this measure caused turmoil in Önancial markets and, most likely, distorted intertemporal choices of households, which is reáected in the large negative innovation to log . The low values of log in the 1990s with respect to the 1970s and 1980s eased the ináationary pressures in the economy.
The shock to the utility of leisure, log , grows in the 1970s and falls in the 1980s to stabilize at a very low value in the 1990s. The likelihood wants to track, in this way, the path of average hours worked: low in the 1970s, increasing in the 1980s, and stabilizing in the 1990s. Higher hours also lower the marginal cost of Örms (wages fall relative to the technology level). The reduction in marginal costs also helped to reduce ináation during Greenspanís tenure.
The evolution of the investment-speciÖc technology, log , shows a sharp drop after 1973 (when it is likely that energy-intensive capital goods su§ered the consequences of the oil shocks in the form of economic obsolescence) and large positive realizations in the late 1990s (our model interprets the sustained boom of those years as the consequence of strong improvements in investment technology). These positive realizations were an additional help to contain ináation during the 1990s. In comparison, the neutral-technology shocks, log , have been stable since 1959, with only a few big shocks at the end of the sample.
The evolution of the monetary policy shock, log , reveals large innovations in the early 1980s. This is due both to the fast change in policy brought about by Volcker and to the fact that a Taylor rule might not fully capture the dynamics of monetary policy during a period in which money growth targeting was attempted. Sims and Zha (2006) also Önd that the Volcker period appears to be one with large disturbances to the policy rule and argue that the Taylor rule formalism can be a misleading perspective from which to view policy during that time. Our evidence from the estimated intertemporal, intratemporal, and investment shocks suggests that monetary authorities faced a more di¢ cult environment in the 1970s and early 1980s than in the 1990s.
As a way to gauge the level of uncertainty of our smoothed estimates, we also plot in Ögure 4.1 the same shock ( 2 standard deviations). The lesson to take away from this Ögure is that, in all cases, the data are informative about the history we just narrated.
Std. Dev. Inter. Shock +/- 2 Std. Dev.

Std. Dev. Intra. Shock +/- 2 Std. Dev.

Std. Dev. Invest. Shock +/- 2 Std. Dev.

Std. Dev. Tech. Shock +/- 2 Std. Dev.

Std. Dev. Mon. Shock +/- 2 Std. Dev.

Figure 4.2: Smoothed standard deviation shocks to the intertemporal shock, the intratemporal shock, the investment-speciÖc shock, the technology shock, and the monetary policy (log t) shock +/- 2 s.d.

We move now, in Ögure 4.2, to plot the evolution of the volatility shocks, all of them in logdeviations with respect to their estimated means (plus/minus two standard deviations). We see in this Ögure that the standard deviation of the intertemporal shock was particularly high in the 1970s and only slowly went down during the 1980s and early 1990s. By the end of the sample, the standard deviation of the intertemporal shock was roughly at the level where it started. In comparison, the standard deviation of all the other shocks is relatively stable except, perhaps, for a large drop in the standard deviation of the monetary policy shock in the early 1980s as well as large changes in the standard deviation of the investment shock during the period of oil price shocks. Hence, the 1970s and the 1980s were more volatile than the 1960s and the 1990s, creating a tougher environment for monetary policy. This result also conÖrms Blanchard and Simonís (2001) and Nason and Smithís (2008) observation that volatility had a downward trend in the 20th century with an abrupt and temporal increase in the 1970s. Also, from the size of the plus/minus two standard deviations, we conclude that the big movements in the di§erent series that we report can be ascertained with a reasonable degree of conÖdence.
Drift on Taylor Rule Param. on Inflation +/- 2 Std. Dev.
Figure 4.3: Smoothed path for the Taylor rule parameter on ináation standard deviations.

Finally, in Ögure 4.3, we plot the evolution of the response of monetary policy to ináation plus/minus a two-standard-deviation interval. In particular, we graph . This graph shows us an intriguing narrative. The parameter started the sample around its estimated mean, slightly over 1, and it grew more or less steadily during the 1960s until reaching a peak in early 1968. After that year, it su§ered a fast collapse that took it below 1 in 1971. To put this evolution in perspective, it is useful to remember that Burns was appointed chairman in February 1970. The parameter stayed below 1 for all of the 1970s. The arrival of Volcker is quickly picked up by our smoothed estimates: it increases to over 2 after a few months and stays high during all the years of Volckerís tenure. Interestingly, our estimate captures well the observation by Goodfriend and King (2007) that monetary policy tightened in the spring of 1980 as ináation and long-run ináation expectations continued to grow. Its level stayed roughly constant at this high during the remainder of Volckerís tenure. But as quickly as it rose when Volcker arrived, it went down again when he departed. Greenspanís tenure at the Fed meant that, by 1990, the response of the monetary authority to ináation was again below 1. Moreover, our estimates are relatively tight. Fern·ndez-Villaverde et al. (2010a) discuss how the results of the estimation relate to historical evidence.
5. Conclusion
In this paper, we have shown how to estimate dynamic equilibrium models with stochastic volatility. The key of the procedure is to realize that a second-order perturbation to the solution of this class of models has a very particular structure that can be easily exploited to build an e¢ cient particle Ölter. The recent boom in the literature on dynamic equilibrium models with stochastic volatility suggests that this procedure may have many uses. Our characterization of the solution might also be, on many occasions, of interest in itself to understand the dynamic properties of the equilibrium even if the researcher does not want to estimate the model.
As an application to illustrate how the procedure works we have estimated a business cycle model with both stochastic volatility in the structural shocks that drive the economy and parameter drifting in the monetary policy rule. Such a model is motivated by the need to have an empirical framework where we can account for the time-varying volatility of U.S. aggregate time series. In particular, we have explained how you get point estimates in such a model and how to recover and analyze the smoothed structural and volatility shocks. Finally, through di§erent comments -even if brief and not exhaustive- we have discussed the di§erent empirical lessons that one can get from all those steps.
References
- [1] Aruoba, S.B., J. Fern·ndez-Villaverde, and J. Rubio-RamÌrez (2006). ìComparing Solution Methods for Dynamic Equilibrium Economies.îJournal of Economic Dynamics and Control 30, 2477-2508.
- [2] Bansal, R. and A. Yaron (2004). ìRisks for the Long Run: a Potential Resolution of Asset Pricing Puzzles.îJournal of Finance 59, 1481-1509.
- [3] Bianchi, F. (2009). ìRegime Switches, AgentsíBeliefs, and Post-World War II U.S. Macroeconomic Dynamics.îMimeo, Duke University.
- [4] Binsbergen, J.V., J. Fern·ndez-Villaverde, R.J. Koijen, and J. Rubio-Ramirez (2012) ìThe Term Structure of Interest Rates in a DSGE Model with Recursive Preferences.î Journal of Monetary Economics, forthcoming.
- [5] Blanchard, O.J. and J. Simon (2001). ìThe Long and Large Decline in U.S. Output Volatility.îBrookings Papers on Economic Activity 2001-1, 187-207.
- [6] Bloom, N., S. Bond, and J. Van Reenen (2007). ìUncertainty and Investment Dynamics. Review of Economic Studies 74, 391-415.
- [7] Bloom, N. (2009). ìThe Impact of Uncertainty Shocks.îEconometrica 77, 623ñ685.
- [8] Bloom, N., N. Jaimovich, N., and M. Floetotto (2012). ìReally Uncertain Business Cycles.î NBER Working Paper 18245.
- [9] Caldara, D., J. Fern·ndez-Villaverde, J. Rubio-Ramirez, and W. Yao (2012). ìComputing DSGE Models with Recursive Preferences and Stochastic Volatility.î Review of Economic Dynamics 15, 188-206.
- [10] Chetty, R., A. Guren, D. Manoli, and A. Weber (2011).ì Micro versus Macro Labor Supply Elasticities.îAmerican Economic Review: Papers and Proceedings, 101, 472-475.
- [11] Clarida, R., J. GalÌ, and M. Gertler (2000). ìMonetary Policy Rules and Macroeconomic Stability: Evidence and Some Theory.îQuarterly Journal of Economics 115, 147-180.
- [12] Doucet A., N. de Freitas and N. Gordon (2001). Sequential Monte Carlo Methods in Practice. Springer Verlag.
- [13] Farmer, R.E., D.F. Waggoner, and T. Zha (2009). ìUnderstanding Markov-Switching Rational Expectations Models.îJournal of Economic Theory 144, 1849-1867.
- [14] Fern·ndez-Villaverde, J. and J. Rubio-RamÌrez (2007). ìEstimating Macroeconomic Models: A Likelihood Approach.îReview of Economic Studies 74, 1059-1087.
- [15] Fern·ndez-Villaverde, J. and J. Rubio-Ramirez (2008). ìHow Structural Are Structural Parameters?îNBER Macroeconomics Annual 2007, 83-137.
- [16] Fern·ndez-Villaverde, J., P. GuerrÛn-Quintana, and J. Rubio-RamÌrez (2009). ìThe New Macroeconometrics: A Bayesian Approach.îin A. OíHagan and M. West (eds.) Handbook of Applied Bayesian Analysis, Oxford University Press.
- [17] Fern·ndez-Villaverde, J., P. GuerrÛn-Quintana, and J. Rubio-RamÌrez (2010a). ìReading the Recent Monetary History of the U.S., 1959-2007î Federal Reserve Bank of St. Louis Review 92(4), 311-38.
- [18] Fern·ndez-Villaverde, J., P. GuerrÛn-Quintana, and J. Rubio-RamÌrez (2010b). ìFortune or Virtue: Time-Variant Volatilities Versus Parameter Drifting in U.S. Data.î NBER Working Papers 15928.
- [19] Fern·ndez-Villaverde, J., P. GuerrÛn-Quintana, J. Rubio-RamÌrez, and M. Uribe (2010c). ìRisk Matters: The Real E§ects of Volatility Shocks.î American Economic Review 101, 2530-61.
- [20] Flury, T. and N. Shephard (2011). ìBayesian Inference Based Only on Simulated Likelihood: Particle Filter Analysis of Dynamic Economic Models.î Econometric Theory 27, 933-956.
- [21] Goodfriend, M. and R. King (2007). ìThe Incredible Volcker Disináation.î Journal of Monetary Economics 52, 981-1015.
- [22] Judd, K. and S. Guu (1997). ìAsymptotic Methods for Aggregate Growth Models.î Journal of Economic Dynamics and Control 21, 1025-1042.
- [23] Justiniano A. and G.E. Primiceri (2008). ìThe Time Varying Volatility of Macroeconomic Fluctuations.îAmerican Economic Review 98, 604-641.
- [24] K¸nsch, H.R. (2005). ìRecursive Monte Carlo Filters: Algorithms and Theoretical Analysis. Annals of Statistics 33, 1983-2021.
- [25] Nason, James M. and Gregor W. Smith (2008). ìGreat Moderation(s) and US Interest Rates: Unconditional Evidence.îThe B.E. Journal of Macroeconomics (Contributions) 8, Art. 30.
- [26] Pitt, M.K. and N. Shephard (1999). ìFiltering via Simulation: Auxiliary Particle Filters. Journal of the American Statistical Association 94, 590-599.
- [27] Santos, M.S. and A. Peralta-Alva (2005). ìAccuracy of Simulations for Stochastic Dynamic Modelsî. Econometrica 73, 1939-1976.
- [28] Schmitt-GrohÈ, S., and M. Uribe (2004). ìSolving Dynamic General Equilibrium Models Using a Second-Order Approximation to the Policy Function.îJournal of Economic Dynamics and Control 28, 755-775.
- [29] Schreft, S. (1990). ìCredit Controls: 1980.î Federal Reserve Bank of Richmond Economic Review 1990, 6, 25-55.
- [30] Shephard, N. (2008). ìStochastic Volatility.î The New Palgrave Dictionary of Economics. Palgrave MacMillan.
- [31] Sims, C.A. and T. Zha (2006). ìWere There Regime Switches in U.S. Monetary Policy?î American Economic Review 96, 54-81.
- [32] Stock, J.H. and M.W. Watson (2002). ìHas the Business Cycle Changed, and Why?îNBER Macroeconomics Annual 17, 159-218.
6. Technical Appendix (Not for Publication)
This technical appendix is organized as follows. First, it presents the proofs of theorems 2 and 3. Second, it describes the equilibrium of the model in more detail. Third, it shows how parameter drifting in the monetary policy rule does not appear in the Örst-order approximation. Fourth, it focuses on describing how we e¢ ciently compute the model. Fifth, it o§ers details on how we elaborated the data. Sixth, it discusses the determinacy of the model. Finally, it includes some additional empirical results not reported in the main text regarding the IRF of the model, the Öt of the model in comparison with some alternatives, and some counterfactuals.
6.1. Proofs
6.1.1. Theorem 2
We start by proving theorem 2. In this theorem, we characterize the Örst-order derivatives of the policy functions h and evaluated at the steady state. We Örst show that the Örst partial derivatives of h and g with respect to any component of , or evaluated at the steady state are zero (in other words, that the Örst-order approximation of the policy functions do not depend on volatility shocks nor their innovations nor on the perturbation parameter). Before proceeding, note that using (2), we can write , in a compact manner, as a function of and
\[\mathcal {Z} _ {t + 1} = \varsigma \left(\mathcal {Z} _ {t}, \Sigma_ {t}, \Lambda \mathcal {E} _ {t + 1}, \Lambda \mathcal {U} _ {t + 1}; \gamma\right),\tag{26}\]
that using can be expressed as a function of , and
\[\Sigma_ {t + 1} = \vartheta \Sigma_ {t} + \eta \Lambda \mathcal {U} _ {t + 1},\tag{27}\]
that using (4) we can write as a function of ; and
\[\mathcal {Z} _ {t} = \varsigma \left(\mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}, \mathcal {E} _ {t}, \mathcal {U} _ {t}; \gamma\right),\tag{28}\]
and that using (5) can be expressed as
\[\Sigma_ {t} = \vartheta \Sigma_ {t - 1} + \eta \mathcal {U} _ {t}\tag{29}\]
where # and are both diagonal matrices with diagonal elements equal to and respectively. If we substitute the policy functions (6)-(8) and (26)-(29) into the set of equilibrium conditions (1), we that
\[\begin{array}{c} F \left(\mathcal {S} _ {t}, \mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}, \mathcal {E} _ {t}, \mathcal {U} _ {t}, \Lambda\right) \equiv \\ \mathbb {E} _ {t} f \left( \begin{array}{c} g \left(h \left(\mathcal {S} _ {t}, \mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}, \mathcal {E} _ {t}, \mathcal {U} _ {t}, \Lambda\right), \varsigma \left(\mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}, \mathcal {E} _ {t}, \mathcal {U} _ {t}\right), \vartheta \Sigma_ {t - 1} + \eta \mathcal {U} _ {t}, \Lambda \mathcal {E} _ {t + 1}, \Lambda \mathcal {U} _ {t + 1}, \Lambda\right), \\ g \left(\mathcal {S} _ {t}, \mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}, \mathcal {E} _ {t}, \mathcal {U} _ {t}, \Lambda\right), h \left(\mathcal {S} _ {t}, \mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}, \mathcal {E} _ {t}, \mathcal {U} _ {t}, \Lambda\right), \mathcal {S} _ {t}, \\ \varsigma \left(\varsigma \left(\mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}, \mathcal {E} _ {t}, \mathcal {U} _ {t}\right), \vartheta \Sigma_ {t - 1} + \eta \mathcal {U} _ {t}, \Lambda \mathcal {E} _ {t + 1}, \Lambda \mathcal {U} _ {t+ 1}\right), \varsigma \left(\mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}, \mathcal {E} _ {t}, \mathcal {U} _ {t}\right) \end{array} \right) = 0 \end{array}\]
where, to ease notation, we do not explicitly write that the functions above depend on .
Proof. We want to show that
\[\left[ h _ {\Sigma_ {t - 1}} \right] _ {j} ^ {i _ {1}} = \left[ g _ {\Sigma_ {t - 1}} \right] _ {j} ^ {i _ {2}} = [ h _ {\mathcal {U} _ {t}} ] _ {j} ^ {i _ {1}} = [ g _ {\mathcal {U} _ {t}} ] _ {j} ^ {i _ {2}} = [ h _ {\Lambda} ] ^ {i _ {1}} = [ g _ {\Lambda} ] ^ {i _ {2}} = 0\]
for , and
We show this result in three steps that basically repeat the same argument based on the homogeneity of a system of linear equations:
1. We write the derivative of the element of F with respect to the element of as
\[\left[ F _ {\Sigma_ {t - 1}} \right] _ {j} ^ {i} = \left[ f _ {\mathcal {Y} _ {t + 1}} \right] _ {i _ {2}} ^ {i} \left(\left[ g _ {\mathcal {S} _ {t}} \right] _ {i _ {1}} ^ {i _ {2}} \left[ h _ {\Sigma_ {t - 1}} \right] _ {j} ^ {i _ {1}} + \left[ g _ {\Sigma_ {t - 1}} \right] _ {j} ^ {i _ {2}} \vartheta_ {j}\right) + \left[ f _ {\mathcal {Y} _ {t}} \right] _ {i _ {2}} ^ {i} \left[ g _ {\Sigma_ {t - 1}} \right] _ {j} ^ {i _ {2}} + \left[ f _ {\mathcal {S} _ {t + 1}} \right] _ {i _ {1}} ^ {i} \left[ h _ {\Sigma_ {t - 1}} \right] _ {j} ^ {i _ {1}} = 0\]
for and . This is a homogeneous system on and for , and . Thus
\[\left[ h _ {\Sigma_ {t - 1}} \right] _ {j} ^ {i _ {1}} = \left[ g _ {\Sigma_ {t - 1}} \right] _ {j} ^ {i _ {2}} = 0\]
for
2. We write the derivative of the element of F with respect to the element of as
\[[ F _ {\mathcal {U} _ {t}} ] _ {j} ^ {i} = \left[ f _ {\mathcal {Y} _ {t + 1}} \right] _ {i _ {2}} ^ {i} \left([ g _ {\mathcal {S} _ {t}} ] _ {i _ {1}} ^ {i _ {2}} [ h _ {\mathcal {U} _ {t}} ] _ {j} ^ {i _ {1}} + \left[ g _ {\Sigma_ {t - 1}} \right] _ {j} ^ {i _ {2}} \left(1 - \vartheta_ {j} ^ {2}\right) ^ {\frac {1}{2}} \eta_ {j}\right) + [ f _ {\mathcal {Y} _ {t}} ] _ {i _ {2}} ^ {i} [ g _ {\mathcal {U} _ {t}} ] _ {j} ^ {i _ {2}} + \left[ f _ {\mathcal {S} _ {t + 1}} \right] _ {i _ {1}} ^ {i} [ h _ {\mathcal {U} _ {t}} ] _ {j} ^ {i _ {1}} = 0\]
for and . Since we have already shown that 0 for and , this is a homogeneous system on and for , and . Thus
\[[ h _ {\mathcal {U} _ {t}} ] _ {j} ^ {i _ {1}} = [ g _ {\mathcal {U} _ {t}} ] _ {j} ^ {i _ {2}} = 0\]
for , and
3. Finally, we write the derivative of the element of F with respect to as
\[[ F _ {\Lambda} ] ^ {i} = \left[ f _ {\mathcal {Y} _ {t + 1}} \right] _ {i _ {2}} ^ {i} \left([ g _ {\mathcal {S} _ {t}} ] _ {i _ {1}} ^ {i _ {2}} [ h _ {\Lambda} ] ^ {i _ {1}} + [ g _ {\Lambda} ] ^ {i _ {2}}\right) + [ f _ {\mathcal {Y} _ {t}} ] _ {i _ {2}} ^ {i} [ g _ {\Lambda} ] ^ {i _ {2}} + \left[ f _ {\mathcal {S} _ {t + 1}} \right] _ {i _ {1}} ^ {i} [ h _ {\Lambda} ] ^ {i _ {1}} = 0\]
for . Since this is a homogeneous system on and for and , we have that
\[\left[ h _ {\Lambda} \right] ^ {i _ {1}} = \left[ g _ {\Lambda} \right] ^ {i _ {2}} = 0\]
for and
6.1.2. Theorem 3
Let us now prove theorem 3. We show, among other things, that the second partial derivatives of h and g with respect to either log and any other variable but are also zero for any . We divide the proof into three parts.
Proof, part 1. The Örst part of the proof deals with the cross-derivatives of the policy functions h and g with respect to and any of , or and it shows that all of them are equal to zero. In particular, we want to show that
\[\left[ h _ {\Lambda , \mathcal {S} _ {t}} \right] _ {j} ^ {i _ {1}} = \left[ g _ {\Lambda , \mathcal {S} _ {t}} \right] _ {j} ^ {i _ {2}} = 0\]
for , and and
\[\left[ h _ {\Lambda , \mathcal {Z} _ {t - 1}} \right] _ {j} ^ {i _ {1}} = \left[ g _ {\Lambda , \mathcal {Z} _ {t - 1}} \right] _ {j} ^ {i _ {2}} = \left[ h _ {\Lambda , \Sigma_ {t - 1}} \right] _ {j} ^ {i _ {1}} = \left[ g _ {\Lambda , \Sigma_ {t - 1}} \right] _ {j} ^ {i _ {2}} = [ h _ {\Lambda , \mathcal {E} _ {t}} ] _ {j} ^ {i _ {1}} = [ g _ {\Lambda , \mathcal {E} _ {t}} ] _ {j} ^ {i _ {2}} = [ h _ {\Lambda , \mathcal {U} _ {t}} ] _ {j} ^ {i _ {1}} = [ g _ {\Lambda , \mathcal {U} _ {t}} ] _ {j} ^ {i _ {2}} = 0\]
for
We show this result in Öve steps. We again exploit the homogeneity of a system of linear equations.
1. We consider the cross-derivative of the element of F with respect to and the element of
\[[ F _ {\Lambda , \mathcal {S} _ {t}} ] _ {j} ^ {i} = \left[ f _ {\mathcal {Y} _ {t + 1}} \right] _ {i _ {2}} ^ {i} \left([ g _ {\mathcal {S} _ {t}} ] _ {i _ {1}} ^ {i _ {2}} [ h _ {\Lambda , \mathcal {S} _ {t}} ] _ {j} ^ {i _ {1}} + [ g _ {\Lambda , \mathcal {S} _ {t}} ] _ {i _ {1}} ^ {i _ {2}} [ h _ {\mathcal {S} _ {t}} ] _ {j} ^ {i _ {1}}\right) + [ f _ {\mathcal {Y} _ {t}} ] _ {i _ {2}} ^ {i} [ g _ {\Lambda , \mathcal {S} _ {t}} ] _ {j} ^ {i _ {2}} + \left[ f _ {\mathcal {S} _ {t + 1}} \right] _ {i _ {1}} ^ {i} [ h _ {\Lambda , \mathcal {S} _ {t}} ] _ {j} ^ {i _ {1}} = 0\]
for and . This is a homogeneous system on and for , and . Thus
\[\left[ h _ {\Lambda , \mathcal {S} _ {t}} \right] _ {j} ^ {i _ {1}} = \left[ g _ {\Lambda , \mathcal {S} _ {t}} \right] _ {j} ^ {i _ {2}} = 0\]
for
2. We consider the cross-derivative of the element of F with respect to and the element of
\[\begin{array}{c} \left[ F _ {\Lambda , \mathcal {Z} _ {t - 1}} \right] _ {j} ^ {i} = \left[ f _ {\mathcal {Y} _ {t + 1}} \right] _ {i _ {2}} ^ {i} \left([ g _ {\mathcal {S} _ {t}} ] _ {i _ {1}} ^ {i _ {2}} \left[ h _ {\Lambda , \mathcal {Z} _ {t - 1}} \right] _ {j} ^ {i _ {1}} + [ g _ {\Lambda , \mathcal {S} _ {t}} ] _ {i _ {1}} ^ {i _ {2}} \left[ h _ {\mathcal {Z} _ {t - 1}} \right] _ {j} ^ {i _ {1}} + \left[ g _ {\Lambda , \mathcal {Z} _ {t - 1}} \right] _ {j} ^ {i _ {2}} \rho_ {j}\right) \\ + [ f _ {\mathcal {Y} _ {t}} ] _ {i _ {2}} ^ {i} \left[ g _ {\Lambda , \mathcal {Z} _ {t - 1}} \right] _ {j} ^ {i _ {2}} + \left[ f _ {\mathcal {S} _ {t + 1}} \right] _ {i _ {1}} ^ {i} \left[ h _ {\Lambda , \mathcal {Z} _ {t - 1}} \right] _ {j} ^ {i _ {1}} = 0 \end{array}\]
for and . Since for and , this is a homogeneous system on and for , , and . Hence
\[\left[ h _ {\Lambda , \mathcal {Z} _ {t - 1}} \right] _ {j} ^ {i _ {1}} = \left[ g _ {\Lambda , \mathcal {Z} _ {t - 1}} \right] _ {j} ^ {i _ {2}} = 0\]
for , and
3. We consider the cross-derivative of the element of F with respect to and the
element of
\[\begin{array}{c} \left[ F _ {\Lambda , \Sigma_ {t - 1}} \right] _ {j} ^ {i} = \left[ f _ {\mathcal {Y} _ {t + 1}} \right] _ {i _ {2}} ^ {i} \left(\left[ g _ {\mathcal {S} _ {t}} \right] _ {i _ {1}} ^ {i _ {2}} \left[ h _ {\Lambda , \Sigma_ {t - 1}} \right] _ {j} ^ {i _ {1}} + \left[ g _ {\Lambda , \Sigma_ {t - 1}} \right] _ {j} ^ {i _ {2}} \vartheta_ {j}\right) \\ + \left[ f _ {\mathcal {Y} _ {t}} \right] _ {i _ {2}} ^ {i} \left[ g _ {\Lambda , \Sigma_ {t - 1}} \right] _ {j} ^ {i _ {2}} + \left[ f _ {\mathcal {S} _ {t + 1}} \right] _ {i _ {1}} ^ {i} \left[ h _ {\Lambda , \Sigma_ {t - 1}} \right] _ {j} ^ {i _ {1}} = 0 \end{array}\]
for and . This is a homogeneous system on and for , and . Hence
\[\left[ h _ {\Lambda , \Sigma_ {t - 1}} \right] _ {j} ^ {i _ {1}} = \left[ g _ {\Lambda , \Sigma_ {t - 1}} \right] _ {j} ^ {i _ {2}} = 0\]
for , and
4. We consider the cross-derivative of the element of F with respect to and the element of
\[\begin{array}{r} [ F _ {\Lambda , \mathcal {E} _ {t}} ] _ {j} ^ {i} = [ f _ {\mathcal {Y} _ {t + 1}} ] _ {i _ {2}} ^ {i} \left([ g _ {\mathcal {S} _ {t}} ] _ {i _ {1}} ^ {i _ {2}} [ h _ {\Lambda , \mathcal {E} _ {t}} ] _ {j} ^ {i _ {1}} + [ g _ {\Lambda , \mathcal {S} _ {t}} ] _ {i _ {1}} ^ {i _ {2}} [ h _ {\mathcal {E} _ {t}} ] _ {j} ^ {i _ {1}} + [ g _ {\Lambda , \mathcal {Z} _ {t - 1}} ] _ {j} ^ {i _ {2}} \sigma_ {j} \exp^ {\vartheta_ {j} \log \sigma_ {j, t - 1}}\right) \\ + [ f _ {\mathcal {Y} _ {t}} ] _ {i _ {2}} ^ {i} [ g _ {\Lambda , \mathcal {E} _ {t}} ] _ {j} ^ {i _ {2}} + [ f _ {\mathcal {S} _ {t + 1}} ] _ {i _ {1}} ^ {i} [ h _ {\Lambda , \mathcal {E} _ {t}} ] _ {j} ^ {i _ {1}} = 0 \end{array}\]
for and . Since for and and for and , this is a homogeneous system on and for , and . Thus
\[\left[ h _ {\Lambda , \mathcal {E} _ {t}} \right] _ {j} ^ {i _ {1}} = \left[ g _ {\Lambda , \mathcal {E} _ {t}} \right] _ {j} ^ {i _ {2}} = 0\]
for , and
5. We consider the cross-derivative of the element of F with respect to and the element of
\[\begin{array}{c} [ F _ {\Lambda , \mathcal {U} _ {t}} ] _ {j} ^ {i} = \left[ f _ {\mathcal {Y} _ {t + 1}} \right] _ {i _ {2}} ^ {i} \left([ g _ {\mathcal {S} _ {t}} ] _ {i _ {1}} ^ {i _ {2}} [ h _ {\Lambda , \mathcal {U} _ {t}} ] _ {j} ^ {i _ {1}} + \left[ g _ {\Lambda , \Sigma_ {t - 1}} \right] _ {j} ^ {i _ {2}} \left(1 - \vartheta_ {j} ^ {2}\right) ^ {\frac {1}{2}} \eta_ {j}\right) \\ + [ f _ {\mathcal {Y} _ {t}} ] _ {i _ {2}} ^ {i} [ g _ {\Lambda , \mathcal {U} _ {t}} ] _ {j} ^ {i _ {2}} + \left[ f _ {\mathcal {S} _ {t + 1}} \right] _ {i _ {1}} ^ {i} [ h _ {\Lambda , \mathcal {U} _ {t}} ] _ {j} ^ {i _ {1}} = 0 \end{array}\]
for and . Since we have shown that for and ; we have that the above system is a homogeneous system on and for , and . Then
\[[ h _ {\Lambda , \mathcal {U} _ {t}} ] _ {j} ^ {i _ {1}} = [ g _ {\Lambda , \mathcal {U} _ {t}} ] _ {j} ^ {i _ {2}} = 0\]
for
Proof, part 2. The second part of the proof deals with the cross-derivatives of the policy functions h and with respect to and any of , or and it shows that all of them are equal to zero with one exception. In particular, we want to show that
\[\left[ h _ {\mathcal {S} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\mathcal {S} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = 0\]
for , and ,
\[\left[ h _ {\mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = \left[ h _ {\Sigma_ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\Sigma_ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = 0\]
for , and , and
\[\left[ h _ {\mathcal {E} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\mathcal {E} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = 0\]
for , and
We show this result in four steps (and where we have already taken advantage of the terms that we know to be equal to zero from previous derivations).
1. We consider the cross-derivative of the element of with respect to the element of and the element of
\[\begin{array}{r} \left[ F _ {\mathcal {S} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i} = \left[ f _ {\mathcal {Y} _ {t + 1}} \right] _ {i _ {2}} ^ {i} \left([ g _ {\mathcal {S} _ {t}} ] _ {i _ {1}} ^ {i _ {2}} \left[ h _ {\mathcal {S} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} + \left[ g _ {\mathcal {S} _ {t}, \Sigma_ {t - 1}} \right] _ {i _ {1}, j _ {2}} ^ {i _ {2}} [ h _ {\mathcal {S} _ {t}} ] _ {j _ {1}} ^ {i _ {1}} \vartheta_ {j _ {2}}\right) \\ + [ f _ {\mathcal {Y} _ {t}} ] _ {i _ {2}} ^ {i} \left[ g _ {\mathcal {S} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} + \left[ f _ {\mathcal {S} _ {t + 1}} \right] _ {i _ {1}} ^ {i} \left[ h _ {\mathcal {S} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = 0 \end{array}\]
for , and . This is a homogeneous system on and 2
and . Therefore
\[\left[ h _ {\mathcal {S} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\mathcal {S} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = 0\]
for
2. We consider the cross-derivative of the element of F with respect to the element of and the element of
\[\begin{array}{c} \left[ F _ {\mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i} \\ = \left[ f _ {\mathcal {Y} _ {t + 1}} \right] _ {i _ {2}} ^ {i} \left([ g _ {\mathcal {S} _ {t}} ] _ {i _ {1}} ^ {i _ {2}} \left[ h _ {\mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} + \left[ g _ {\mathcal {S} _ {t}, \Sigma_ {t - 1}} \right] _ {i _ {1}, j _ {2}} ^ {i _ {2}} \left[ h _ {\mathcal {Z} _ {t - 1}} \right] _ {j _ {1}} ^ {i _ {1}} \vartheta_ {j _ {2}} + \left[ g _ {\mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} \rho_ {j _ {1}} \vartheta_ {j _ {2}}\right) \\ + [ f _ {\mathcal {Y} _ {t}} ] _ {i _ {2}} ^ {i} \left[ g _ {\mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} + \left[ f _ {\mathcal {S} _ {t + 1}} \right] _ {i _ {1}} ^ {i} \left[ h _ {\mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = 0 \end{array}\]
for , and . Since we just found that 0 for , and , this is a homogeneous system on and for , and . Therefore
\[\left[ h _ {\mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = 0\]
for
3. We consider the cross-derivative of the element of F with respect to the element of and the element of
\[\begin{array}{c} \left[ F _ {\Sigma_ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i} = \\ \left[ f _ {\mathcal {Y} _ {t + 1}} \right] _ {i _ {2}} ^ {i} \left(\left[ g _ {\mathcal {S} _ {t}} \right] _ {i _ {1}} ^ {i _ {2}} \left[ h _ {\Sigma_ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} + \left[ g _ {\Sigma_ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} \vartheta_ {j _ {1}} \vartheta_ {j _ {2}}\right) \\ + \left[ f _ {\mathcal {Y} _ {t}} \right] _ {i _ {2}} ^ {i} \left[ g _ {\Sigma_ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} + \left[ f _ {\mathcal {S} _ {t + 1}} \right] _ {i _ {1}} ^ {i} \left[ h _ {\Sigma_ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = 0 \end{array}\]
for and . This is a homogeneous system on and for , and therefore
\[\left[ h _ {\Sigma_ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\Sigma_ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = 0\]
for , and
4. We consider the cross-derivative of the element of with respect to the element of and the element of
\[\begin{array}{c} {\left[ F _ {{\mathcal {E}} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i} =} \\ {\left[ f _ {{\mathcal {Y}} _ {t + 1}} \right] _ {i _ {2}} ^ {i} \left([ g _ {{\mathcal {S}} _ {t}} ] _ {i _ {1}} ^ {i _ {2}} \left[ h _ {{\mathcal {E}} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} + \left[ g _ {{\mathcal {S}} _ {t}, \Sigma_ {t - 1}} \right] _ {i _ {1}, j _ {2}} ^ {i _ {2}} [ h _ {{\mathcal {E}} _ {t}} ] _ {j _ {1}} ^ {i _ {1}} \vartheta_ {j _ {2}} + \left[ g _ {{\mathcal {Z}} _ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} \sigma_ {j _ {1}} \exp^ {\vartheta_ {j _ {1}} \log \sigma_ {j _ {1}, t - 1}} \vartheta_ {j _ {2}}\right)} \\ + [ f _ {{\mathcal {Y}} _ {t}} ] _ {i _ {2}} ^ {i} \left[ g _ {{\mathcal {E}} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} + \left[ f _ {{\mathcal {S}} _ {t + 1}} \right] _ {i _ {1}} ^ {i} \left[ h _ {{\mathcal {E}} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = 0 \end{array}\]
for and . Since we know that for , and , this is a homogeneous system on and for , and if . Therefore
\[\left[ h _ {\mathcal {E} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\mathcal {E} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = 0\]
for , and
Note that if , we have that
\[\begin{array}{c} \left[ F _ {\mathcal {E} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {1}} ^ {i} = \left[ f _ {\mathcal {Y} _ {t + 1}} \right] _ {i _ {2}} ^ {i} * \\ * \left( \begin{array}{c} [ g _ {\mathcal {S} _ {t}} ] _ {i _ {1}} ^ {i _ {2}} \left[ h _ {\mathcal {E} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {1}} ^ {i _ {1}} + \left[ g _ {\mathcal {S} _ {t}, \Sigma_ {t - 1}} \right] _ {i _ {1}, j _ {1}} ^ {i _ {2}} [ h _ {\mathcal {E} _ {t}} ] _ {j _ {1}} ^ {i _ {1}} \vartheta_ {j _ {1}} + \\ \left(\left[ g _ {\mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {1}} ^ {i _ {2}} + \left[ g _ {\mathcal {Z} _ {t - 1}} \right] _ {j _ {1}} ^ {i _ {2}}\right) \sigma_ {j _ {1}} \exp^ {\vartheta_ {j _ {1}} \log \sigma_ {j _ {1}, t - 1}} \vartheta_ {j _ {1}} \\ + [ f _ {\mathcal {Y} _ {t}} ] _ {i _ {2}} ^ {i} \left[ g _ {\mathcal {E} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {1}} ^ {i _ {2}} + \left[ f _ {\mathcal {S} _ {t + 1}} \right] _ {i _ {1}} ^ {i} \left[ h _ {\mathcal {E} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {1}} ^ {i _ {1}} \\ + \left([ f _ {\mathcal {Z} _ {t}} ] _ {j _ {1}} ^ {i} + \left[ f _ {\mathcal {Z} _ {t + 1}} \right] _ {j _ {1}} ^ {i} \rho_ {j _ {1}}\right) \sigma_ {j _ {1}} \exp^ {\vartheta_ {j _ {1}} \log \sigma_ {j _ {1}, t - 1}} \vartheta_ {j _ {1}} = 0 \end{array} \right. \end{array}\]
and since and are di§erent from zero in general for and ; we have that this system is not homogeneous and
\[\left[ h _ {\mathcal {E} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {1}} ^ {i _ {1}} = \left[ g _ {\mathcal {E} _ {t}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {1}} ^ {i _ {2}} \neq 0\]
for
Proof, part 3. The Önal part of the proof deals with the cross-derivatives of the policy functions h and g with respect to and any of , or and it shows that all of them are equal to zero with one exception. In particular, we want to show that
\[\left[ h _ {\mathcal {S} _ {t}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\mathcal {S} _ {t}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = 0\]
for , and ,
\[\left[ h _ {\mathcal {Z} _ {t - 1}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\mathcal {Z} _ {t - 1}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = \left[ h _ {\Sigma_ {t - 1}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\Sigma_ {t - 1}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = [ h _ {\mathcal {U} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = [ g _ {\mathcal {U} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = 0\]
for , and , and
\[\left[ h _ {\mathcal {E} _ {t}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\mathcal {E} _ {t}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = 0\]
for , and
Again, we follow the same steps for each part of the result as before and use our previous Öndings regarding which terms are zero.
1. We consider the cross derivative of the element of F with respect to the element of and the element of
\[\begin{array}{r} \left[ F _ {\mathcal {S} _ {t}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i} = \left[ f _ {\mathcal {Y} _ {t + 1}} \right] _ {i _ {2}} ^ {i} \left(\left[ g _ {\mathcal {S} _ {t}} \right] _ {i _ {1}} ^ {i _ {2}} \left[ h _ {\mathcal {S} _ {t}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} + \left[ g _ {\mathcal {S} _ {t}, \Sigma_ {t - 1}} \right] _ {i _ {1}, j _ {2}} ^ {i _ {2}} \left[ h _ {\mathcal {S} _ {t}} \right] _ {j _ {1}} ^ {i _ {1}} \left(1 - \vartheta_ {j _ {2}} ^ {2}\right) ^ {\frac {1}{2}} \eta_ {j _ {2}}\right) \\ + \left[ f _ {\mathcal {Y} _ {t}} \right] _ {i _ {2}} ^ {i} \left[ g _ {\mathcal {S} _ {t}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} + \left[ f _ {\mathcal {S} _ {t + 1}} \right] _ {i _ {1}} ^ {i} \left[ h _ {\mathcal {S} _ {t}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = 0 \end{array}\]
for , and . Since for , and , this is a homogeneous system on and .Therefore
\[\left[ h _ {\mathcal {S} _ {t}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\mathcal {S} _ {t}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = 0\]
for
2. We consider the cross-derivative of the element of F with respect to the element
of and the element of
\[\begin{array}{c} {\left[ F _ {\mathcal {Z} _ {t - 1}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i} =} \\ {\left[ f _ {\mathcal {Y} _ {t + 1}} \right] _ {i _ {2}} ^ {i} \left( \begin{array}{c} {[ g _ {\mathcal {S} _ {t}} ] _ {i _ {1}} ^ {i _ {2}} \left[ h _ {\mathcal {Z} _ {t - 1}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} + \left[ g _ {\mathcal {S} _ {t}, \Sigma_ {t - 1}} \right] _ {i _ {1}, j _ {2}} ^ {i _ {2}} \left[ h _ {\mathcal {Z} _ {t}} \right] _ {j _ {1}} ^ {i _ {1}} \left(1 - \vartheta_ {j _ {2}} ^ {2}\right) ^ {\frac {1}{2}} \eta_ {j _ {2}}} \\ + \left[ g _ {\mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} \rho_ {j _ {1}} \left(1 - \vartheta_ {j _ {2}} ^ {2}\right) ^ {\frac {1}{2}} \eta_ {j _ {2}} \end{array} \right)} \\ + [ f _ {\mathcal {Y} _ {t}} ] _ {i _ {2}} ^ {i} \left[ g _ {\mathcal {Z} _ {t - 1}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} + \left[ f _ {\mathcal {S} _ {t + 1}} \right] _ {i _ {1}} ^ {i} \left[ h _ {\mathcal {Z} _ {t - 1}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = 0 \end{array}\]
for , and . Since for , and , this is a homogeneous system on and for , and Therefore
\[\left[ h _ {\mathcal {Z} _ {t - 1}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\mathcal {Z} _ {t - 1}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = 0\]
for
3. We consider the cross-derivative of the element of with respect to the element of and the element of
\[\begin{array}{c} \left[ F _ {\Sigma_ {t - 1}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i} = \\ \left[ f _ {\mathcal {Y} _ {t + 1}} \right] _ {i _ {2}} ^ {i} \left(\left[ g _ {\mathcal {S} _ {t}} \right] _ {i _ {1}} ^ {i _ {2}} \left[ h _ {\Sigma_ {t - 1}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} + \left[ g _ {\Sigma_ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} \vartheta_ {j _ {1}} \left(1 - \vartheta_ {j _ {2}} ^ {2}\right) ^ {\frac {1}{2}} \eta_ {j _ {2}}\right) \\ + \left[ f _ {\mathcal {Y} _ {t}} \right] _ {i _ {2}} ^ {i} \left[ g _ {\Sigma_ {t - 1}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} + \left[ f _ {\mathcal {S} _ {t + 1}} \right] _ {i _ {1}} ^ {i} \left[ h _ {\Sigma_ {t - 1}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = 0 \end{array}\]
for , and . Since for , and , this is a homogeneous system on and for , and . Therefore
\[\left[ h _ {\Sigma_ {t - 1}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = \left[ g _ {\Sigma_ {t - 1}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = 0\]
for
4. We consider the cross-derivative of the element of with respect to the element
of and the element of
\[\begin{array}{c} {[ F _ {\mathcal {U} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {2}} ^ {i} =} \\ {\left[ f _ {\mathcal {Y} _ {t + 1}} \right] _ {i _ {2}} ^ {i} \left(\left[ g _ {\Sigma_ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} \left(1 - \vartheta_ {j _ {1}} ^ {2}\right) ^ {\frac {1}{2}} \eta_ {j _ {1}} \left(1 - \vartheta_ {j _ {2}} ^ {2}\right) ^ {\frac {1}{2}} \eta_ {j _ {2}} + [ g _ {\mathcal {S} _ {t}} ] _ {i _ {1}} ^ {i _ {2}} [ h _ {\mathcal {U} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {2}} ^ {i _ {1}}\right)} \\ + [ f _ {\mathcal {Y} _ {t}} ] _ {i _ {2}} ^ {i} [ g _ {\mathcal {U} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {2}} ^ {i _ {2}} + \left[ f _ {\mathcal {S} _ {t + 1}} \right] _ {i _ {1}} ^ {i} [ h _ {\mathcal {U} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = 0 \end{array}\]
for and . Since for and , this is a homogeneous system on and for , and . Therefore
\[[ h _ {\mathcal {U} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = [ g _ {\mathcal {U} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = 0\]
for
5. Finally, consider the cross-derivative of the element of with respect to the element of and the element of
\[\begin{array}{c} {[ F _ {\mathcal {E} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {2}} ^ {i} =} \\ {\left[ f _ {\mathcal {Y} _ {t + 1}} \right] _ {i _ {2}} ^ {i} \left( \begin{array}{c} {[ g _ {\mathcal {S} _ {t}} ] _ {i _ {1}} ^ {i _ {2}} [ h _ {\mathcal {E} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {2}} ^ {i _ {1}} + \left[ g _ {\mathcal {S} _ {t}, \Sigma_ {t - 1}} \right] _ {i _ {1}, j _ {2}} ^ {i _ {2}} [ h _ {\mathcal {E} _ {t}} ] _ {j _ {1}} ^ {i _ {1}} \left(1 - \vartheta_ {j _ {2}} ^ {2}\right) ^ {\frac {1}{2}} \eta_ {j _ {2}}} \\ + \left[ g _ {\mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} \sigma_ {j _ {1}} \exp^ {\vartheta_ {j _ {1}} \log \sigma_ {j _ {1}, t - 1}} \left(1 - \vartheta_ {j _ {2}} ^ {2}\right) ^ {\frac {1}{2}} \eta_ {j _ {2}}} \end{array} \right) \\ + [ f _ {\mathcal {Y} _ {t}} ] _ {i _ {2}} ^ {i} [ g _ {\mathcal {E} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {2}} ^ {i _ {2}} + \left[ f _ {\mathcal {S} _ {t + 1}} \right] _ {i _ {1}} ^ {i} [ h _ {\mathcal {E} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = 0 \end{array}\]
for and . Since for , and , this is a homogeneous system on and for , and if . Therefore
\[\left[ h _ {\mathcal {E} _ {t}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {2}} = \left[ g _ {\mathcal {E} _ {t}, \mathcal {U} _ {t}} \right] _ {j _ {1}, j _ {2}} ^ {i _ {1}} = 0\]
for , and
Note that if , we have that
\[\begin{array}{c} [ F _ {\mathcal {E} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {1}} ^ {i} = \\ \left[ f _ {\mathcal {Y} _ {t + 1}} \right] _ {i _ {2}} ^ {i} \left( \begin{array}{c} \left(\left[ g _ {\mathcal {Z} _ {t - 1}, \Sigma_ {t - 1}} \right] _ {j _ {1}, j _ {1}} ^ {i _ {2}} + \left[ g _ {\mathcal {Z} _ {t - 1}} \right] _ {j _ {1}} ^ {i _ {2}}\right) \sigma_ {j _ {1}} \exp^ {\vartheta_ {j _ {1}} \log \sigma_ {j _ {1}, t - 1}} \left(1 - \vartheta_ {j _ {1}} ^ {2}\right) ^ {\frac {1}{2}} \eta_ {j _ {1}} \\ + \left[ g _ {\mathcal {S} _ {t}, \Sigma_ {t - 1}} \right] _ {i _ {1}, j _ {1}} ^ {i _ {2}} [ h _ {\mathcal {E} _ {t}} ] _ {j _ {1}} ^ {i _ {1}} \left(1 - \vartheta_ {j _ {1}} ^ {2}\right) ^ {\frac {1}{2}} \eta_ {j _ {1}} + [ g _ {\mathcal {S} _ {t}} ] _ {i _ {1}} ^ {i _ {2}} [ h _ {\mathcal {E} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {1}} ^ {i _ {1}} \\ + [ f _ {\mathcal {Y} _ {t}} ] _ {i _ {2}} ^ {i} [ g _ {\mathcal {E} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {1}} ^ {i _ {2}} + \left[ f _ {\mathcal {S} _ {t + 1}} \right] _ {i _ {1}} ^ {i} [ h _ {\mathcal {E} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {1}} ^ {i _ {1}} \\ + \left([ f _ {\mathcal {Z} _ {t}} ] _ {j _ {1}} ^ {i} + \rho_ {j _ {1}} \left[ f _ {\mathcal {Z} _ {t + 1}} \right] _ {j _ {1}} ^ {i}\right) \sigma_ {j _ {1}} \exp^ {\vartheta_ {j _ {1}} \log \sigma_ {j _ {1}, t - 1}} \left(1 - \vartheta_ {j _ {1}} ^ {2}\right) ^ {\frac {1}{2}} \eta_ {j _ {1}} = 0 \end{array} \right. \end{array}\]
and since and are di§erent from zero in general for and ; we have that this system is not homogeneous and hence
\[[ h _ {\mathcal {E} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {1}} ^ {i _ {2}} = [ g _ {\mathcal {E} _ {t}, \mathcal {U} _ {t}} ] _ {j _ {1}, j _ {1}} ^ {i _ {1}} \neq 0\]
for , and ■
6.2. Equilibrium
In this section we describe the equilibrium conditions of the model. First, we introduce the ones related to the household, then the ones related to the Örm and the monetary authority, and Önally we present the market clearing and aggregation conditions.
6.2.1. Households
We can deÖne two Lagrangian multipliers, , the multiplier associated with the budget constraint, and (the marginal Tobinís Q), the multiplier associated with the investment adjustment constraint normalized by . Thus, the Örst-order conditions of the household problem with respect to , and can be written as:
\[d _ {t} \left(c _ {j t} - h c _ {j t - 1}\right) ^ {- 1} - b \beta \mathbb {E} _ {t} d _ {t + 1} \left(c _ {j t + 1} - h c _ {j t}\right) ^ {- 1} = \lambda_ {j t},\tag{30}\]
\[\lambda_ {j t} = \beta \mathbb {E} _ {t} \{\lambda_ {j t + 1} \frac {R _ {t}}{\Pi_ {t + 1}} \},\tag{31}\]
\[r _ {t} = \mu_ {t} ^ {- 1} \Phi^ {\prime} [ u _ {j t} ],\tag{32}\]
\[q _ {j t} = \beta \mathbb {E} _ {t} \left\{\frac {\lambda_ {j t + 1}}{\lambda_ {j t}} \left((1 - \delta) q _ {j t + 1} + r _ {t + 1} u _ {j t + 1} - \mu_ {t + 1} ^ {- 1} \Phi [ u _ {j t + 1} ]\right) \right\},\tag{33}\]
and
\[1 = q _ {j t} \mu_ {t} \left(1 - V \left[ \frac {x _ {j t}}{x _ {j t - 1}} \right] - V ^ {\prime} \left[ \frac {x _ {j t}}{x _ {j t - 1}} \right] \frac {x _ {j t}}{x _ {j t - 1}}\right) + \beta \mathbb {E} q _ {j t + 1} \mu_ {t + 1} \frac {\lambda_ {j t + 1}}{\lambda_ {j t}} V ^ {\prime} \left[ \frac {x _ {j t + 1}}{x _ {j t}} \right] \left(\frac {x _ {j t + 1}}{x _ {j t}}\right) ^ {2}.\tag{34}\]
The Örst-order conditions of the ìlabor packerîimply a demand function for labor:
\[l _ {j t} = \left(\frac {w _ {j t}}{w _ {t}}\right) ^ {- \eta} l _ {t} ^ {d} \quad \forall j\]
and, together with a zero proÖt condition , an expression for the wage:
\[w _ {t} = \left(\int_ {0} ^ {1} w _ {j t} ^ {1 - \eta} d j\right) ^ {\frac {1}{1 - \eta}}.\]
Households follow a Calvo pricing mechanism when they set their wages. At the start of every period, a randomly selected fraction of households can reoptimize their wages. All other households simply index their nominal wages given past ináation with an indexation parameter
Since we postulated in the main text both complete Önancial markets for the households and separable utility in consumption, the marginal utilities of consumption are the same for all households. Thus, in equilibrium, , and
The last two equalities are the most relevant to simplify our analysis: they tell us that the shadow cost of consumption is equated across households and that all households that can reset their wages optimally will do it at the same level . With these two results, and after several steps of algebra, we Önd that the evolution of wages is described by two recursive equations:
\[f _ {t} = \frac {\eta - 1}{\eta} (w _ {t} ^ {*}) ^ {1 - \eta} \lambda_ {t} w _ {t} ^ {\eta} l _ {t} ^ {d} + \beta \theta_ {w} \mathbb {E} _ {t} \left(\frac {\Pi_ {t} ^ {\chi_ {w}}}{\Pi_ {t + 1}}\right) ^ {1 - \eta} \left(\frac {w _ {t + 1} ^ {*}}{w _ {t} ^ {*}}\right) ^ {\eta - 1} f _ {t + 1}\tag{35}\]
and
\[f _ {t} = \psi d _ {t} \varphi_ {t} \left(\frac {w _ {t}}{w _ {t} ^ {*}}\right) ^ {\eta (1 + \vartheta)} \left(l _ {t} ^ {d}\right) ^ {1 + \vartheta} + \beta \theta_ {w} \mathbb {E} _ {t} \left(\frac {\prod_ {t} ^ {\chi_ {w}}}{\prod_ {t + 1}}\right) ^ {- \eta (1 + \vartheta)} \left(\frac {w _ {t + 1} ^ {*}}{w _ {t} ^ {*}}\right) ^ {\eta (1 + \vartheta)} f _ {t + 1}\tag{36}\]
on the auxiliary variable .
Taking advantage of the expression for the wage and that, in every period, a fraction of households set as their wage and the remaining fraction partially index their nominal wage by past ináation, we can write the law of motion of real wage as:
\[w _ {t} ^ {1 - \eta} = \theta_ {w} \left(\frac {\prod_ {t - 1} ^ {\chi_ {w}}}{\prod_ {t}}\right) ^ {1 - \eta} w _ {t - 1} ^ {1 - \eta} + (1 - \theta_ {w}) w _ {t} ^ {* 1 - \eta}.\tag{37}\]
6.2.2. Firms
The Önal good producer is perfectly competitive and minimizes its costs subject to the production function (24) and taking as given all intermediate goods prices and the Önal good price The optimality conditions of this problem result in a demand function for each intermediate good with the classic form:
\[y _ {i t} = \left(\frac {p _ {i t}}{p _ {t}}\right) ^ {- \varepsilon} y _ {t} ^ {d} \quad \forall i\]
where is the aggregate demand and a price for the Önal good:
\[p _ {t} = \left(\int_ {0} ^ {1} p _ {i t} ^ {1 - \varepsilon} d i\right) ^ {\frac {1}{1 - \varepsilon}}.\]
Each of the intermediate goods is produced by a monopolistic competitor. Intermediate good producers produce the quantity demanded of the good by renting and at prices and . Then, by minimization, we have a marginal cost of:
\[m c _ {t} = \left(\frac {1}{1 - \alpha}\right) ^ {1 - \alpha} \left(\frac {1}{\alpha}\right) ^ {\alpha} \frac {w _ {t} ^ {1 - \alpha} r _ {t} ^ {\alpha}}{A _ {t}}\tag{38}\]
The marginal cost is constant for all Örms and all production levels given , and .
Given the demand function, the intermediate good producers set prices to maximize proÖts. However, when they do so, they follow the same Calvo pricing scheme as households. In each period, a fraction of intermediate good producers reoptimize their prices. All other Örms partially index their prices by past ináation with an indexation parameter
The solution for the Örmís pricing problem has a recursive structure in two new auxiliary variables and that take the form:
\[g _ {t} ^ {1} = \lambda_ {t} m c _ {t} y _ {t} ^ {d} + \beta \theta_ {p} \mathbb {E} _ {t} \left(\frac {\Pi_ {t} ^ {\chi}}{\Pi_ {t + 1}}\right) ^ {- \varepsilon} g _ {t + 1} ^ {1},\tag{39}\]
\[g _ {t} ^ {2} = \lambda_ {t} \Pi_ {t} ^ {*} y _ {t} ^ {d} + \beta \theta_ {p} \mathbb {E} _ {t} \left(\frac {\Pi_ {t} ^ {\chi}}{\Pi_ {t + 1}}\right) ^ {1 - \varepsilon} \left(\frac {\Pi_ {t} ^ {*}}{\Pi_ {t + 1} ^ {*}}\right) g _ {t + 1} ^ {2},\tag{40}\]
and
\[\varepsilon g _ {t} ^ {1} = (\varepsilon - 1) g _ {t} ^ {2}\tag{41}\]
where
\[\Pi_ {t} ^ {*} = \frac {p _ {t} ^ {*}}{p _ {t}}\tag{42}\]
is the ratio between the optimal new price (common across all Örms that can reset their prices) and the price of the Önal good. With this structure, the price index follows:
\[p _ {t} ^ {1 - \varepsilon} = \theta_ {p} \left(\Pi_ {t - 1} ^ {\chi}\right) ^ {1 - \varepsilon} p _ {t - 1} ^ {1 - \varepsilon} + (1 - \theta_ {p}) p _ {t} ^ {* 1 - \varepsilon}\]
or, normalizing by :
\[1 = \theta_ {p} \left(\frac {\Pi_ {t - 1} ^ {\chi}}{\Pi_ {t}}\right) ^ {1 - \varepsilon} + (1 - \theta_ {p}) \Pi_ {t} ^ {* 1 - \varepsilon}.\tag{43}\]
6.2.3. The Monetary Authority
The model is closed with a monetary authority that sets the nominal interest rates by a modiÖed Taylor rule described in (25).
6.2.4. Market Clearing and Aggregation
Aggregate demand is given by:
\[y _ {t} ^ {d} = c _ {t} + x _ {t} + \mu_ {t} ^ {- 1} \Phi [ u _ {t} ] k _ {t - 1}.\tag{44}\]
By relying on the observation that the capital-labor ratio is constant across Örms, we can derive that aggregate supply is:
\[y _ {t} ^ {s} = \frac {A _ {t} (u _ {t} k _ {t - 1}) ^ {\alpha} (l _ {t} ^ {d}) ^ {1 - \alpha} - \phi z _ {t}}{v _ {t} ^ {p}}\tag{45}\]
where:
\[v _ {t} ^ {p} = \int_ {0} ^ {1} \left(\frac {p _ {i t}}{p _ {t}}\right) ^ {- \varepsilon} d i\]
is the aggregate loss of e¢ ciency induced by price dispersion of the intermediate goods.
Market clearing requires that
\[y _ {t} = y _ {t} ^ {d} = y _ {t} ^ {s}.\tag{46}\]
By the properties of Calvoís pricing:
\[v _ {t} ^ {p} = \theta_ {p} \left(\frac {\Pi_ {t - 1} ^ {\chi}}{\Pi_ {t}}\right) ^ {- \varepsilon} v _ {t - 1} ^ {p} + (1 - \theta_ {p}) \Pi_ {t} ^ {* - \varepsilon}.\tag{47}\]
Finally, demanded labor is given by:
\[l _ {t} ^ {d} = \frac {1}{\upsilon_ {t} ^ {w}} \int_ {0} ^ {1} l _ {j t} d j = l _ {t}\tag{48}\]
where:
\[\upsilon_ {t} ^ {w} = \int_ {0} ^ {1} \left(\frac {w _ {j t}}{w _ {t}}\right) ^ {- \eta} d j.\]
is the aggregate loss of labor input induced by wage dispersion among di§erentiated types of labor. Again, by Calvoís pricing, this ine¢ ciency evolves as:
\[\upsilon_ {t} ^ {w} = \theta_ {w} \left(\frac {w _ {t - 1}}{w _ {t}} \frac {\Pi_ {t - 1} ^ {\chi_ {w}}}{\Pi_ {t}}\right) ^ {- \eta} \upsilon_ {t - 1} ^ {w} + (1 - \theta_ {w}) (\Pi_ {t} ^ {w *}) ^ {- \eta}.\tag{49}\]
Thus an equilibrium is characterized by equations (30)-(49), the Taylor rule (25), the law of motion for the structural shocks (augmented with the parameter drifts), and the law of motion for the volatility shocks.
6.3. Non-linearities in Parameter Drifting
We argued in the main text that, since we wanted to consider the e§ects of stochastic volatility on our model, it was of the essence to deal with higher-order approximations. In this section we argue that higher-order approximations are also key to dealing with parameter drifting in the Taylor rule. The reason is that parameter drifting disappears from a linear solution. To see this, take the Taylor rule deÖned in (25) (assuming only in this paragraph and to simplify notation that and log and let us rewrite it:
\[f \left(\widehat {R} _ {t}, \widehat {R} _ {t - 1}, \widehat {\Pi} _ {t}, \widehat {\gamma} _ {\Pi , t}, \varepsilon_ {\xi t}\right) = \exp^ {\widehat {R} _ {t}} - \exp^ {\gamma_ {R} \widehat {R} _ {t - 1} + (1 - \gamma_ {R}) \gamma_ {\Pi} \exp^ {\widehat {\gamma} _ {\Pi , t}} \widehat {\Pi} _ {t} + \sigma_ {\xi} \varepsilon_ {\xi t}} = 0.\]
where we have expressed each variable in terms of log deviation with respect to the steady state, . The log-linear approximation of f around the steady state is
\[f \left(\widehat {R} _ {t}, \widehat {R} _ {t - 1}, \widehat {\Pi} _ {t}, \widehat {\gamma} _ {\Pi , t}, \varepsilon_ {\xi t}\right) \simeq f _ {1} (\mathbf {0}) \widehat {R} _ {t} + f _ {2} (\mathbf {0}) \widehat {R} _ {t - 1} + f _ {3} (\mathbf {0}) \widehat {\Pi} _ {t} + f _ {4} (\mathbf {0}) \widehat {\gamma} _ {\Pi , t} + f _ {5} (\mathbf {0}) \varepsilon_ {\xi t}.\]
where is the Örst derivative of the function f evaluated at (0; 0; 0; 0; 0) with respect to the variable i. Note that:
\[f _ {4} \left(\widehat {R} _ {t}, \widehat {R} _ {t - 1}, \widehat {\Pi} _ {t}, \widehat {\gamma} _ {\Pi , t}, \varepsilon_ {\xi t}\right) = - \left(1 - \gamma_ {R}\right) \gamma_ {\Pi} \widehat {\Pi} _ {t} \exp^ {\widehat {\gamma} _ {\Pi , t}} \exp^ {\gamma_ {R} \widehat {R} _ {t - 1} + (1 - \gamma_ {R}) \gamma_ {\Pi} \exp^ {\widehat {\gamma} _ {\Pi , t}} \widehat {\Pi} _ {t} + \sigma_ {\xi} \varepsilon_ {\xi t}}.\]
Clearly, and does not play any role in this Örst-order approximation. This is the consequence of one variable (t) being raised to another variable . Thus, the log-linear approximation of the Taylor rule is:
\[f \left(\widehat {R} _ {t}, \widehat {R} _ {t - 1}, \widehat {\Pi} _ {t}, \widehat {\gamma} _ {\Pi , t}, \varepsilon_ {m t}\right) \simeq \widehat {R} _ {t} - \gamma_ {R} \widehat {R} _ {t - 1} - (1 - \gamma_ {R}) \gamma_ {\Pi} \widehat {\Pi} _ {t} + \sigma_ {\xi} \varepsilon_ {\xi t}\tag{50}\]
which does not depend on ; but only on the steady state and it is exactly the same expression as the one without parameter drifting. Hence, in order to capture parameter drifting in the Taylor rule, we need, at least, to perform a second-order approximation.
6.4. Computation
In this section, we provide some more details regarding the computation of the paper. We generate all the derivatives required by our second-order perturbation with Mathematica 6.0. In that way, we do not need to recompute the derivatives, the most time-intensive step, for each set of parameter values in our estimation. Once we have all the relevant derivatives, we export them automatically into Fortran Öles. This whole process takes about 3 hours. Then, we compile the resulting Öles with the Intel Fortran Compiler version 10:1:025 with IMSL. Previous versions failed to compile our project because of the length of some of the expressions. Compilation takes about 18 hours. The project has 1798 Öles and occupies 2:33 Gbytes of memory.
The next step is, for given parameter values, to compute the Örst- and second-order approximation to the decision rules around the deterministic steady state using the analytic derivatives we found before. For this task, Fortran takes around 5 seconds. Once we have the solution, we approximate the likelihood using the particle Ölter with 10; 000 particles. This number delivered a good compromise between accuracy and time to compute the likelihood. The evaluation of one likelihood requires 22 seconds on a Dell server with 8 processors. Once we have the likelihood evaluation, we guess new parameter values and we start again. This means that drawing 5,000 times from the posterior (even forgetting about the initial search over a grid of parameter values) takes around 38 hours.
It is important to emphasize that the Mathematica and Fortran code were highly optimized in order to 1) keep the size of the project within reasonable dimensions (otherwise, the compiler cannot parse the Öles and, even when it can, it delivers code that is too ine¢ cient) and 2) provide a fast computation of the likelihood.
Perhaps the most important task in that optimization was the parallelization of the Fortran code using OPENMP as well as the compilation options: OG (global optimizations) and Loop Unroll. In addition, we tailored specialized code to perform the matrix multiplications required in the Örstand second-order terms of our model solution.
Implementing corollary 1 requires the solution of a linear system of equations and the computation of a Jacobian. For our particular application, we found that the following sequence of LAPACK operations delivered the fastest solution:
1. DGESV (computes the solution to a real system of linear equations A X = B).
2. DGETRI (computes the inverse of a matrix using the LU factorization from the previous line).
3. DGETRF (helps to compute the determinant of the inverse from the previous line).
Without the parallelization and our optimized code, the solution of the model and evaluation of its likelihood take about 70 seconds.
With respect to the random-walk Metropolis-Hastings, we performed an intensive process of Öne-tuning of the chain, both in terms of initial conditions as well as in terms of getting the right acceptance level. The only other important remark is to remember that as pointed out by McFadden (1989) and Pakes and Pollard (1989), we must keep the random numbers used for resampling in the particle Ölter constant across draws of the Markov chain. This is required to achieve stochastic equi-continuity, and even if this condition is not strictly necessary in a Bayesian framework, it reduces the numerical variance of the procedure, which was a serious concern for us given the complexity of our problem.
6.5. Construction of Data
When we estimate the model, we make the series provided by the National Income and Product Accounts (NIPA) consistent with the deÖnition of variables in the theory. The main adjustment we undertake is to express both real output and real gross investment in consumption units. Our model implies that there is a numeraire in terms of which all the other prices need to be quoted. We pick consumption as the numeraire. The NIPA, in comparison, uses an index of all prices to transform nominal GDP and investment into real values. In the presence of changing relative prices, such as the ones we have seen in the U.S. over the last several decades with the fall in the relative price of capital, NIPAís procedure biases the valuation of di§erent series in real terms.
We map theory into the data by computing our own series of real output and real investment. To do so, we use the relative price of investment, deÖned as the ratio of an investment deáator and a deáator for consumption. The denominator is easily derived from the deáators of nondurable goods and services reported in the NIPA. It is more complicated to obtain the numerator because, historically, NIPA investment deáators were poorly constructed. Instead, we rely on the investment deáator computed by Fisher (2006). Since the series ends early in 2000.Q4, we have extended it to 2007.Q1 by following Fisherís methodology.
For the real output per capita series, we Örst deÖne nominal output as nominal consumption plus nominal gross investment. We deÖne nominal consumption as the sum of personal consumption expenditures on non-durable goods and services. We deÖne nominal gross investment as the sum of personal consumption expenditures on durable goods, private residential investment, and non-residential Öxed investment. Per capita nominal output is equal to the ratio between our nominal output series and the civilian non-institutional population between 16 and 65. To obtain per capita values, we divide the previous series by the civilian non-institutional population between 16 and 65. Finally, real wages are deÖned as compensation per hour in the non-farm business sector divided by the CPI deáator.
6.6. Determinacy of Equilibrium
We mentioned in the main text that the estimated value of (1.045 in levels) guarantees local determinacy of the equilibrium. To see this note that, for local determinacy, the relevant part of the solution of the model is only the linear, Örst-order component. This component depends on , the mean policy response, and not on the current value of . The economic intuition is that local unicity survives even if temporarily violates the Taylor principle as long as there is reversion to the mean in the policy response and, thus, the agents have the expectation that will satisfy the Taylor principle on average. For a related result in models with Markov-switching regime changes, see Davig and Leeper (2006). While we cannot Önd an analytical expression for the determinacy region, numerical experiments show that, conditional on the other point estimates, values of above 0.98 ensure uniqueness. Since the likelihood assigns zero probability to values of lower than 1.01, well inside the determinacy region, multiplicity of local equilibria is not an issue in our application.
6.7. Impulse Response Functions
As a check of our estimates, we plot the IRFs generated by the model to a monetary policy shock. This exercise is an important test. If the IRFs match the shapes and sizes of those gathered by time series methods such as SVARs, it will strengthen our belief in the rest of our results. Otherwise, we should at least understand where the di§erences come from.
Auspiciously, the answer is positive: our model generates dynamic responses that are close to the ones from SVARs (see, for instance, Sims and Zha, 2006). The left panel of Ögure 6.1 plots the IRFs to a monetary shock of three variables commonly discussed in monetary models: the federal funds rate, output growth, and ináation. Since we have a non-linear model, in all the Ögures in this section, we report the generalized IRFs starting from the mean of the ergodic distribution (Koop, Pesaran, and Potter, 1996). After a one-standard-deviation shock to the federal funds rate, ináation goes down in a hump-shaped pattern for many quarters and output growth drops.
The right panel of Ögure 6.1 plots the IRFs after a one-standard-deviation innovation to the monetary policy shock computed conditional on Öxing to the estimated mean during the tenure of each of three di§erent chairmen of the Board of Governors: the combination Burns-Miller, Volcker, and Greenspan. This exercise tells us how the variation on the systematic component of monetary policy has a§ected the dynamics of aggregate variables. Furthermore, it allows a comparison with numerous similar exercises done in the literature with SVARs where the IRFs are estimated on di§erent subsamples.
The most interesting di§erence is that the response of output growth under Volcker was the mildest: the estimated average stance of monetary policy under his tenure reduces the volatility of output. Ináation responds moderately as well since the agents have the expectation that future shocks will be smoothed out by the monetary authority. This Önding also explains why the IRFs of the interest rate are nearly on top of each other for all three periods: while we estimate that monetary policy responded more during Volckerís years for any given level of ináation than under Burns-Miller or Greenspan, this policy lowers ináation deviations and hence moderates the actual movement along the equilibrium path of the economy. Moreover, this second set of IRFs already points out one important result of this paper: we estimate that monetary policy under Burns-Miller and Greenspan was similar, while it was di§erent under Volcker. This Önding will be reinforced by the results we present in the main body of the text.
IRF Monetary Policy Shock

Conditional IRF Monetary Policy Shock Figure 6.1: IRFs to a Monetary Policy Shock, Unconditional and Conditional

For completeness, we also plot, in Ögure 6.2, the IRFs to each of the other four shocks in our model: the two preferences shocks (intertemporal and intratemporal) and the two technology shocks (investment-speciÖc and neutral). The behavior of the model is standard. A one-standarddeviation intertemporal preference shock raises output growth and ináation because there is an increase in the desire for consumption in the current period. The intratemporal shock lowers output because labor becomes less attractive, driving up the marginal costs and, with it, prices. The two supply shocks raise output growth and lower ináation by increasing productivity. All of those IRFs show that the behavior of the model is standard.



Figure 6.2: IRFs of ináation, output growth, and the federal funds rate to an intertemporal demand shock, an intratemporal demand shock, an investment-speciÖc shock, and a neutral technology shock. The responses are measured as log di§erences with respect to the mean of the ergodic distribution.

6.8. Model Comparison
Another use of our procedure to evaluate the likelihood function is to compare our model against alternative models -or alternative versions of the same model. For instance, a natural question is to compare our benchmark model with stochastic volatility and parameter drifting with a version without parameter drifting but with stochastic volatility. That is: once we have included stochastic volatility, is it still important to allow for changes in monetary policy to account for the time-varying volatility of aggregate data in the U.S.?
Given our Bayesian framework, a natural approach for model comparison is the computation of log marginal data densities (log MDD) and log Bayes factors. Let us focus on the proposed example of comparing the full model with stochastic volatility and parameter drifting with a version without parameter drifting (no drift) but with stochastic volatility. In this second case, we have two fewer parameters, and (but we still have . To ease notation, we partition the parameter vector as ; where is the vector of all the other parameters, common to the two versions of the model.
Given that our priors are 1) uniform, 2) independent of each other, and 3) cover all the areas where the likelihood is (numerically) positive, and that 4) the priors on are common across the two speciÖcations of the model, we can write
\[\log p \left(\mathbb {Y} ^ {T} = \mathbb {Y} ^ {\text {data}, T}; d r i f t\right) = \log \int p \left(\mathbb {Y} ^ {T} = \mathbb {Y} ^ {\text {data}, T}; \gamma , d r i f t\right) d \gamma + \log p (\widetilde {\gamma}) + \log p \left(\rho_ {\gamma_ {\Pi}}\right) + \log p \left(\sigma_ {\pi}\right),\]
where log , log , and log are constants and
\[\log p \left(\mathbb {Y} ^ {T} = \mathbb {Y} ^ {\text { data }, T}; n o \text { drift }\right) = \log \int p \left(\mathbb {Y} ^ {T} = \mathbb {Y} ^ {\text { data }, T}; \widetilde {\gamma}, n o \text { drift }\right) d \widetilde {\gamma} + \log p (\widetilde {\gamma}).\]
Thus
\[\begin{array}{r l} & {\log B _ {d r i f t, n o d r i f t} = \log p \left(\mathbb {Y} ^ {T} = \mathbb {Y} ^ {d a t a, T}; d r i f t\right) - \log p \left(\mathbb {Y} ^ {T} = \mathbb {Y} ^ {d a t a, T}; n o d r i f t\right)} \\ & {= \log \int p \left(\mathbb {Y} ^ {T} = \mathbb {Y} ^ {d a t a, T}; \gamma , d r i f t\right) d \widetilde {\gamma} - \log \int p \left(\mathbb {Y} ^ {T} = \mathbb {Y} ^ {d a t a, T}; \widetilde {\gamma}, n o d r i f t\right) d \widetilde {\gamma}} \\ & {\qquad - \log p \left(\rho_ {\gamma_ {\Pi}}\right) - \log p \left(\sigma_ {\pi}\right).} \end{array}\]
The di§erence
\[\log \int p \left(\mathbb {Y} ^ {T} = \mathbb {Y} ^ {d a t a, T}; \gamma , d r i f t\right) d \gamma - \log \int p \left(\mathbb {Y} ^ {T} = \mathbb {Y} ^ {d a t a, T}; \widetilde {\gamma}, n o d r i f t\right) d \widetilde {\gamma}\]
tells us how much better the version with parameter drift Öts the data in comparison with the version with no drift. The last two terms, log ; penalize for the presence of two extra parameters.
We estimate the log MDDs following Gewekeís (1998) harmonic mean method. This requires us to generate a new draw of the posterior of the model for the speciÖcation with no parameter drift to compute log ; no drift d . After doing so, we Önd that
\[\log B _ {d r i f t, n o d r i f t} = 1 2 6. 1 3 3 1 + \log p (\rho_ {\gamma_ {\Pi}}) + \log p (\sigma_ {\pi})\]
This expression shows a potential problem of Bayes factors: by picking uniform priors for and spread out over a su¢ ciently large interval, we could overcome any di§erence in Öt. But the prior for is pinned down by our desire to keep that process stationary, which imposes natural bounds in [ 1; 1] and makes log . Thus, there is only one degree of freedom left: our choice of log . Any sensible prior for will only put mass in a relatively small interval: the point estimate is 0:1479, the standard deviation is 0:002, and the likelihood is numerically zero for values bigger than 0.2. Hence, we can safely impose that log would imply a uniform prior between 0 and 2:7183, a considerably wider support than any evidence in the data), and conclude that log . This is conventionally considered overwhelming evidence in favor of the model with parameter drift (Je§reys, 1961, for instance, suggests that di§erences bigger than 5 are decisive). Thus, even after controlling for stochastic volatility, the data strongly prefer a speciÖcation where monetary policy has changed over time. This Önding, however, does not imply that volatility shocks did not play an important role in the time-varying volatilities of U.S. aggregate time series. In fact, as we will see in the next section of this appendix, they were a key mechanism in accounting for it.4
It has been noted that the estimation of log MDDs is dangerous because of numerical instabilities in the evaluation of the integral log marginal data density (log MDD). This concern is particularly relevant in our case, since we have a large model saddled with burdensome computation. Thus, as a robustness analysis, we also computed the Bayesian Information Criterion (BIC) (Schwarz, 1978). The BIC, which avoids the need to handle the integral in the log MDD, can be understood as an asymptotic approximation of the Bayes factor that also automatically penalizes for extra parameters. The BIC of model i is deÖned:
\[B I C _ {i} = - 2 \ln p \left(\mathbb {Y} ^ {T} = \mathbb {Y} ^ {d a t a, T}; \widehat {\gamma}, i\right) + k _ {i} \ln n\]
where is the maximum likelihood estimator (or, given our áat priors, the mode of the posterior), is the number of parameters, and n is the number of observations. Then, the BIC of the model with stochastic volatility and parameter drifting is If we eliminate parameter drifting and the parameters and associated with it (and, of course, with a new point estimate of the other parameters) 7; 484:7. The di§erence is, therefore, of over 138 log points, which is again overwhelming evidence in favor of the model with parameter drifting.
4A formal comparison with the case without stochastic volatility is more di¢ cult, since we are taking advantage of its presence to evaluate the likelihood. Fortunately, Justiniano and Primiceri (2008) and Fern·ndez-Villaverde and Rubio-RamÌrez (2007) estimate models similar to ours with and without stochastic volatility (in the Örst case, using only a Örst-order approximation to the decision rules of the agents and in the second with measurement errors). Both papers Önd that the Öt of the model improves substantially when we include stochastic volatility. Finally, Fern·ndez-Villaverde and Rubio-RamÌrez (2008) compare a model with parameter drifting and no stochastic volatility with a model without parameter drifting and no stochastic volatility and report that parameter drifting is also strongly preferred by the likelihood.
6.9. Historical Counterfactuals
One important additional exercise is to quantify how much of the observed changes in the volatility of aggregate U.S. variables can be accounted for by changes in the standard deviations of shocks and how much by changes in policy. To accomplish this, we build a number of historical counterfactuals. In these exercises, we remove one source of variation at a time and we measure how aggregate variables would have behaved when hit only by the remaining shocks. Since our model is structural in the sense of Hurwicz (1962) (it is invariant to interventions, including shocks by nature such as the ones we are postulating), we will obtain an answer that is robust to the Lucas critique.
In the next two subsections, we will always plot the same three basic variables that we used in section 4 of the main text: ináation, output growth, and the federal funds rate. Counterfactual histories of other variables could be built analogously. Also, we will have vertical bars for the tenure of each chairman, following the same color scheme as in section 4.
6.9.1. Counterfactual I: Switching Chairmen
In our Örst counterfactual, we move one chairman from his mandate to an alternative time period. For example, we appoint Greenspan as chairman during the Burns-Miller years. By that, we mean that the Fed would have followed the policy rule dictated by the average estimated during Greenspanís time while starting from the same states as Burns-Miller and su§ering the same shocks (both structural and of volatility). We repeat this exercise with all the other possible combinations: Volcker in the Burns-Miller decade, Burns-Miller in Volckerís mandate, Greenspan in Volckerís time, Burns-Miller in the Greenspan years, and, Önally, Volcker in Greenspanís time.
It is important to be careful in interpreting this exercise. By appointing Greenspan at Volckerís time, we do not literally mean Greenspan as a person, but Greenspan as a convenient label for a particular monetary policy response to shocks that according to our model were observed during his tenure. The real Greenspan could have behaved in a di§erent way, for example, as a result of some non-linearities in monetary policy that are not properly captured by a simple rule such as the one postulated in section 3. The argument could be pushed one step further and we could think about the appointment of Volcker as an endogenous response of the political-economic equilibrium to high ináation. In our model agents have a probability distribution regarding possible changes in monetary policy in the next periods, but those changes are uncorrelated with current conditions. Therefore, our model cannot capture the endogeneity of policy selection.
Another issue that we sidelined is the evolution of expectations. In our model, agents have rational expectations and observe the changes in monetary policy parameters. This hypothesis may be a poor approximation of the agentsíbehavior in real life. It could be the case that was high in 1984, even though ináation was already low by that time, because of the high ináationary expectations that economic agents held during most of the 1980s (this point is also linked to issues of commitment and credibility that our model does not address). While we see all these arguments as interesting lines of research, we Önd it important to focus Örst on our basic counterfactual.
Moments In table 6.1, we report the mean and the standard deviation of ináation, output growth, and the federal funds rate in the observed and in the sets of counterfactual data. Ináation was high with Burns-Miller, fell with Volcker, and stayed low with Greenspan. Output growth went down during the Volcker years to recover with Greenspan. The federal funds rate reached its peak with Volcker. The standard deviation of output growth fell from 4.7 in Burns-Millerís time to 2.45 with Greenspan, a cut in half. Similarly, ináation volatility fell nearly 54 percent and the federal funds rate volatility 5 percent.
But table 6.1 also tells us one important result: time-varying monetary policy signiÖcantly a§ected average ináation. In particular, Volckerís response to ináation was strong and switching him to either Burns-Millerís or Greenspanís time would have reduced average ináation dramatically. But it also tells us other things: contrary to the conventional wisdom, our estimates suggest that the stance of monetary policy against ináation under Greenspan was not strong. In Burns-Millerís time, the monetary policy under Greenspan would have delivered slightly higher average ináation, 6.83 versus the observed 6.23, accompanied by a lower federal funds rate and lower output growth, 1.89 versus the observed 2.03. The di§erence is even bigger in Volckerís time, during which average ináation would have been nearly 1.4 percent higher, while output growth would have been virtually identical (1.34 versus 1.38). The key for this Önding is in the behavior of the federal funds rate, which would have increased only by 9 basis points, on average, if Greenspan had been in charge of the Fed instead of Volcker. Given the higher ináation in the counterfactual, the higher nominal interest rates would have meant much lower real rates. The counterfactual of Burns-Miller in Greenspanís and Volckerís time casts doubt on the malignant reputations of these two short-lived chairmen, at least when compared with Greenspan. Burns-Miller would have brought even slightly lower ináation than Greenspan, thanks to a higher real federal funds rate and a bit higher output growth. However, Burns-Miller would have delivered higher ináation than Volcker.
Table 6.1: Switching Chairmen, Data versus Counterfactual Histories
| Means | Standard Deviations | |||||
| Inflation | Output Gr. | FFR | Inflation | Output Gr. | FFR | |
| BM (data) | 6.2333 | 2.0322 | 6.5764 | 2.7347 | 4.7010 | 2.2720 |
| Greenspan to BM | 6.8269 | 1.8881 | 6.5046 | 3.3732 | 4.6781 | 2.0103 |
| Volcker to BM | 4.3604 | 1.5010 | 7.6479 | 2.4620 | 4.6219 | 2.3470 |
| Volcker (data) | 5.3584 | 1.3846 | 10.3338 | 3.1811 | 4.4811 | 3.4995 |
| BM to Volcker | 6.4132 | 1.3560 | 10.4126 | 2.9728 | 4.4220 | 3.0648 |
| Greenspan to Volcker | 6.7284 | 1.3423 | 10.4235 | 2.9824 | 4.3730 | 2.8734 |
| Greenspan (data) | 2.9583 | 1.5177 | 4.7352 | 1.2675 | 2.4567 | 2.1887 |
| BM to Greenspan | 2.3355 | 1.5277 | 4.4529 | 1.5625 | 2.4684 | 2.4652 |
| Volcker to Greenspan | -0.4947 | 1.3751 | 3.6560 | 1.7700 | 2.4705 | 2.7619 |
This is an important empirical Önding: according to our model, time-varying monetary policy signiÖcantly a§ected average ináation. While Volckerís response to ináation was strong, Greenspanís response was milder and he seems to have behaved quite similarly to how Burns-Miller would have behaved.
Counterfactual Paths An alternative way to analyze our results is to plot the whole counterfactual histories summarized in table 6.1. We Önd it interesting to plot the whole history because changes in the economyís behavior in one period will propagate over time and we want to understand, for example, how Greenspanís legacy would have molded Volckerís tenure. Also, plotting the whole history allows us to track the counterfactual response of monetary policy to large economic events such as the oil shocks.
In Ögure 6.3, we move to Burns-Miller being reappointed in Greenspanís time. This plot suggests that the di§erences in monetary policy under Greenspan and Burns-Miller may have been overstated by the literature.


Figure 6.3: Burns-Miller during the Greenspan years

In Ögure 6.4, we plot the counterfactual of Burns-Miller extending their tenure to 1987. The results are very similar to the case in which we move Greenspan to the same period: slower disináation and no improvement in output growth.

Figure 6.4: Burns-Miller during the Volcker years


A particularly interesting exercise is to check what would have happened if Reagan had decided to reappoint Volcker and not appoint Greenspan. We plot these results in Ögure 6.5. The quick answer is: lower ináation and interest rates. Our estimates also suggest that Volcker would have reduced price increases with little cost to output.

Figure 6.5: Volcker during the Greenspan years


Our Önal exercise is to plot, in Ögure 6.6, the counterfactual in which we move Volcker to the time of Burns-Miller. The main Önding is that ináation would have been rather lower, especially because the e§ects of the second oil shock would have been much more muted. This counterfactual is plausible: other countries, such as Germany, Switzerland, and Japan, that undertook a more aggressive monetary policy during the late 1970s were able to keep ináation under control at levels below 5 percent at an annual rate, while the U.S. had peaks of price increases over 10 percent.


Figure 6.6: Volcker during the Burns-Miller years

6.9.2. Counterfactual II: No Volatility Changes
In our second historical counterfactual, we compute how the economy would have performed in the absence of changes in the volatility of the shocks, that is, if the volatility of the innovation of the structural shocks had been Öxed at its historical mean. To do so, we back out the smoothed structural shocks as we did in section 4.7 and we feed them to the model, given our parameter point estimates and the historical mean of volatility, to generate series for ináation, output, and the federal funds rate.
Moments Table 6.2 reports the moments of the data (in annualized terms) and the moments from the counterfactual history (no s.v. in the table stands for ìno stochastic volatilityî). In both cases, we include the moments for the whole sample and for the sample divided before and after
1984.Q1, a conventional date for the start of the great moderation (McConnell and PÈrez-QuirÛs, 2000). In the last two rows of the table, we compute the ratio of the moments after 1984.Q1 over the moments before 1984.Q1. The benchmark model with stochastic volatility plus parameter drifting replicates the data exactly.
Some of the numbers in table 6.2 are well known. For instance, after 1984, the standard deviation of ináation falls by nearly 60 percent, the standard deviation of output growth falls by 44 percent, and the standard deviation of the federal funds rate falls by 39 percent. In terms of means, after 1984, there is less ináation and the federal funds rate is lower, but output growth is also 15 percent lower.
Table 6.2: No Volatility Changes, Data versus Counterfactual History
| Means | Standard Deviations | |||||
| Inflation | Output Growth | FFR | Inflation | Output Growth | FFR | |
| Data | 3.8170 | 1.8475 | 6.0021 | 2.6181 | 3.5879 | 3.3004 |
| Data, pre 1984.1 | 4.6180 | 1.9943 | 6.7179 | 3.2260 | 4.3995 | 3.8665 |
| Data, after 1984.1 | 2.9644 | 1.6911 | 5.2401 | 1.3113 | 2.4616 | 2.3560 |
| No s.v. | 2.5995 | 0.7169 | 6.9388 | 3.5534 | 3.1735 | 2.4128 |
| No s.v., pre-1984.1 | 2.0515 | 0.9539 | 6.3076 | 3.7365 | 3.4120 | 2.7538 |
| No s.v., after-1984.1 | 3.1828 | 0.4647 | 7.6106 | 3.2672 | 2.8954 | 1.7673 |
| Data, post-1984.1/pre-1984.1 | 0.6419 | 0.8480 | 0.7800 | 0.4065 | 0.5595 | 0.6093 |
| No s.v., post-1984.1/pre-1984.1 | 1.5515 | 0.4871 | 1.2066 | 0.8744 | 0.8486 | 0.6418 |
The table also reáects the fact that without volatility shocks, the reduction in volatility observed after 1984 would have been noticeably smaller. The standard deviation of ináation would have fallen by only 13 percent, the standard deviation of output growth would have fallen by 16 percent, and the standard deviation of the federal funds rate would have fallen by 35 percent, that is, only 33, 20, and 87 percent, respectively, of how much they would have fallen otherwise. We must resist here the temptation to undertake a standard variance decomposition exercise. Since we have a second-order approximation to the policy function and its associated cross-product terms, we cannot neatly divide total variance among the di§erent shocks as we could do in the linear case.
Table 6.2 documents that, while time-varying policy is reáected in changes of average ináation, volatility shocks a§ect the standard deviation of the observed series. Without time-varying volatility the decrease in observed volatility would not have been nearly as big as we observed in the data. Hence, while changes in the systematic component of monetary policy account for changes in average ináation, volatility shocks account for changes in the standard deviation of ináation, output growth and interest rates observed after 1984. Also, without stochastic volatility, output growth would have been quite lower on average.
Counterfactual Paths Figure 6.7 compares the whole path of the counterfactual history (blue line) with the observed one (red line). Figure 6.7 tells us that volatility shocks mattered throughout the sample. The run-up in ináation would have been much slower in the late 1960s (ináation would have actually been negative during the last years of Martinís tenure) with small e§ects on output growth or the federal funds rate (except at the very end of the sample). Ináation would not have picked up nearly as much during the Örst oil shock, but output growth would have su§ered. During Volckerís time, ináation would also have fallen faster with little cost to output growth. These are indications that both Burns-Miller and Volcker su§ered from large and volatile shocks to the economy. In comparison, during the 1990s, ináation would have been more volatile, with a big increase in the middle of the decade. Similarly, during those years, output growth would have been much lower, with a long recession between 1994 and 1998, and the federal funds rate would have been prominently higher. ConÖrming the results presented in section 4 of the paper, this is another manifestation of how placid the 1990s were for policy makers.


Figure 6.7: Counterfactual ìNo Changes in Volatilityî.

References
- [1] Davig, T. and E.M. Leeper (2006). ìGeneralizing the Taylor Principle.î American Economic Review 97, 607-635.
- [2] Fern·ndez-Villaverde, J. and J. Rubio-RamÌrez (2007). ìEstimating Macroeconomic Models: A Likelihood Approach.îReview of Economic Studies 74, 1059-1087.
- [3] Fernandez-Villaverde, J. and J. Rubio-Ramirez (2008). ìHow Structural Are Structural Parameters?îNBER Macroeconomics Annual 2007, 83-137.
- [4] Fisher, J., (2006). ìThe Dynamic E§ects of Neutral and Investment-SpeciÖc Technology Shocks.îJournal of Political Economy 114, 413-52.
- [5] Geweke, J. (1998). ìUsing Simulation Methods for Bayesian Econometric Models: Inference, Development and Communication.îSta§ Report 249, Federal Reserve Bank of Minneapolis.
- [6] Hurwicz, L. (1962). ìOn the Structural Form of Interdependent Systems.îLogic, Methodology and Philosophy of Science, Proceedings of the 1960 International Congress, 232-39.
- [7] Je§reys H. (1961). The Theory of Probability (3rd ed.). Oxford University Press.
- [8] Justiniano A. and G.E. Primiceri (2008). ìThe Time Varying Volatility of Macroeconomic Fluctuations.îAmerican Economic Review 98, 604-641.
- [9] Koop G., M. Hashem Pesaran, and S. Potter (1996). ìImpulse Response Analysis in Nonlinear Multivariate Models.îJournal of Econometrics 74, 199-147.
- [10] McConnell, M.M. and G. PÈrez-QuirÛs (2000). ìOutput Fluctuations in the United States: What Has Changed Since the Early 1980ís?îAmerican Economic Review 90, 1464-1476.
- [11] McFadden, D.L. (1989). ìA Method of Simulated Moments for Estimation of Discrete Response Models Without Numerical Integration.îEconometrica 57, 995-1026.
- [12] Pakes, A. and D. Pollard (1989). ìSimulation and the Asymptotics of Optimization Estimators.îEconometrica 57, 1027-1057.
- [13] Schwarz, G.E. (1978). ìEstimating the Dimension of a Model.î Annals of Statistics 6, 461ñ 464.
- [14] Sims, C.A. and T. Zha (2006). ìWere There Regime Switches in U.S. Monetary Policy? American Economic Review 96, 54-81.
- 2013-23: “Estimating Dynamic Equilibrium Models with Stochastic Volatility”, Jesús Fernández-Villaverde, Pablo Guerrón-Quintana y Juan F. Rubio-Ramírez.
- 2013-22: “Perturbation Methods for Markov-Switching DSGE Models”, Andrew Foerster, Juan Rubio-Ramirez, Dan Waggoner y Tao Zha.
- 2013-21: “Do Spanish informal caregivers come to the rescue of dependent people with formal care unmet needs?”, Sergi Jiménez-Martín y Cristina Vilaplana Prieto.
- 2013-20: “When Credit Dries Up: Job Losses in the Great Recession”, Samuel Bentolila, Marcel Jansen, Gabriel Jiménez y Sonia Ruano.
- 2013-19: “Efectos de género en las escuelas, un enfoque basado en cohortes de edad”, Antonio Ciccone y Walter Garcia-Fontes.
- 2013-18: “Oil Price Shocks, Income, and Democracy“, Markus Brückner , Antonio Ciccone y Andrea Tesei.
- 2013-17: “Rainfall Risk and Religious Membership in the Late Nineteenth-Century US”, Philipp Ager y Antonio Ciccone.
- 2013-16: “Immigration in Europe: Trends, Policies and Empirical Evidence”, Sara de la Rica, Albrecht Glitz y Francesc Ortega.
- 2013-15: “The impact of family-friendly policies on the labor market: Evidence from Spain and Austria”, Sara de la Rica y Lucía Gorjón García.
- 2013-14: “Gender Gaps in Performance Pay: New Evidence from Spain”, Sara de la Rica, Juan J. Dolado y Raquel Vegas.
- 2013-13: “On Gender Gaps and Self-Fulfilling Expectation: Alternative Implications of Paid-For Training”, Juan J. Dolado, Cecilia García-Peñalosa y Sara de la Rica.
- 2013-12: “Financial incentives, health and retirement in Spain”, Pilar García‐Gómez, Sergi Jiménez‐Martín y Judit Vall Castelló.
- 2013-11: “Gender quotas and the quality of politicians”, Audinga Baltrunaite, Piera Bello, Alessandra Casarico y Paola Profeta.
- 2013-10: “Brechas de Género en los Resultados de PISA :El Impacto de las Normas Sociales y la Transmisión Intergeneracional de las Actitudes de Género”, Sara de la Rica y Ainara González de San Román.
- 2013-09: “¿Cómo escogen los padres la escuela de sus hijos? Teoría y evidencia para España”, Caterina Calsamiglia, Maia Güell.
- 2013-08: “Evaluación de un programa de educación bilingüe en España: El impacto más allá del aprendizaje del idioma extranjero”, Brindusa Anghel, Antonio Cabrales y Jesús M. Carro.
- 2013-07: “Publicación de los resultados de las pruebas estandarizadas externas: ¿Tiene ello un efecto sobre los resultados escolares?”, Brindusa Anghel, Antonio Cabrales, Jorge Sainz e Ismael Sanz.
- 2013-06: “DYPES: A Microsimulation model for the Spanish retirement pension system”, F. J. Fernández-Díaz, C. Patxot y G. Souto.
- 2013-05: “Vertical differentiation, schedule delay and entry deterrence: Low cost vs. full service airlines”, Jorge Validoa, M. Pilar Socorroa y Francesca Medda.
- 2013-04: “Dropout Trends and Educational Reforms: The Role of the LOGSE in Spain”, Florentino Felgueroso, María Gutiérrez‐Domènech y Sergi Jiménez‐Martín.
- 2013-03: “Understanding Different Migrant Selection Patterns in Rural and Urban Mexico”, Simone Bertoli, Herbert Brücker y Jesús Fernández-Huertas Moraga.
- 2013-02: “Understanding Different Migrant Selection Patterns in Rural and Urban Mexico”, Jesús Fernández-Huertas Moraga.
- 2013-01: “Publicizing the results of standardized external tests: Does it have an effect on school outcomes?, Brindusa Anghel, Antonio Cabrales, Jorge Sainz y Ismael Sanz.
- 2012-12: “Visa Policies, Networks and the Cliff at the Border”, Simone Bertoli, Jesús Fernández-Huertas Moraga.
- 2012-11: “Intergenerational and Socioeconomic Gradients of Child Obesity”, Joan Costa-Fonta y Joan Gil.
- 2012-10: “Subsidies for resident passengers in air transport markets”, Jorge Valido, M. Pilar Socorro, Aday Hernández y Ofelia Betancor.
- 2012-09: “Dual Labour Markets and the Tenure Distribution: Reducing Severance Pay or Introducing a Single Contract?”, J. Ignacio García Pérez y Victoria Osuna.
- 2012-08: “The Influence of BMI, Obesity and Overweight on Medical Costs: A Panel Data Approach”, Toni Mora, Joan Gil y Antoni Sicras-Mainar.
- 2012-07: “Strategic behavior in regressions: an experimental”, Javier Perote, Juan Perote-Peña y Marc Vorsatz.
- 2012-06: “Access pricing, infrastructure investment and intermodal competition”, Ginés de Rus y M. Pilar Socorro.
- 2012-05: “Trade-offs between environmental regulation and market competition: airlines, emission trading systems and entry deterrence”, Cristina Barbot, Ofelia Betancor, M. Pilar Socorro y M. Fernanda Viecens.