‹ Volver a la ficha Doc. dt2018-13

Documento de Trabajo - 2018/13 Inference in Bayesian Proxy-SVARs

Jonas E. Arias (Federal Reserve Bank of Philadelphia)

Juan F. Rubio-Ramírez (Emory University, Federal Reserve Bank of Atlanta, and BBVA Research)

Daniel F. Waggoner (Federal Reserve Bank of Atlanta)

fedea

Las opiniones recogidas en este documento son las de sus autores y no coinciden necesariamente con las de FEDEA.

Jonas E. Arias Juan F. Rubio-Ramírez Daniel F. Waggoner

November 16, 2018

Abstract

Motivated by the increasing use of external instruments to identify structural vector autoregressions (SVARs), we develop algorithms for exact finite sample inference in this class of time series models, commonly known as proxy-SVARs. Our algorithms make independent draws from the normal-generalized-normal family of conjugate posterior distributions over the structural parameterization of a proxy-SVAR. Importantly, our techniques can handle the case of set identification and hence they can be used to relax the additional exclusion restrictions unrelated to the external instruments often imposed to facilitate inference when more than one instrument is used to identify more than one equation as in Mertens and Montiel-Olea (2018). JEL classification: C15; C32

Keywords: SVARs; External Instruments; Importance Sampler

The views expressed in this paper are solely those of the authors and do not necessarily reflect the views of the Federal Reserve Bank of Atlanta, the Federal Reserve Bank of Philadelphia, or the Federal Reserve System. Any errors or omissions are the responsibility of the authors. No statements here should be treated as legal advice.
∗Federal Reserve Bank of Philadelphia Email: jonas.arias@phil.frb.org
†Emory University, Federal Reserve Bank of Atlanta, and BBVA Research Email: juan.rubio-ramirez@emory.edu
‡Federal Reserve Bank of Atlanta. Email: daniel.f.waggoner@atl.frb.org

1 Introduction

The method of identification of structural vector autoregressions (SVARs) with external instruments, commonly known as proxy-SVARs, has grown to become influential in empirical macroeconomics. Currently, most of the papers using proxy-SVARs work under the frequentist paradigm.1 But, while substantial progress has been made on such front, less is known about conducting Bayesian inference in this class of structural time series models; exceptions are Bahaj (2014), Drautzburg (2016), and Caldara and Herbst (2016).

In this paper we contribute to this line of research by developing eficient algorithms to independently draw from the family of restricted normal-generalized-normal posterior distributions over the structural parameterization of a proxy-SVAR conditional on exogeneity restrictions. These restrictions require that the correlation between the proxies and some of the structural shocks be equal to zero. The fact that we can draw independently opens the door to use the Bayesian paradigm in larger models.

Our main algorithm combines the sampler developed by Waggoner and Zha (2003) with a variant of the importance sampler developed by Arias, Rubio-Ramírez, and Waggoner (2018). The fundamental insight in the latter was to produce independent draws from a normal-inverse-Wishart distribution over the reduced-form parameters and generalize the QR decomposition to produce independent draws from a distribution over the orthogonal-reduced-form parameterization of the SVAR conditional on sign and zero restrictions. These draws were then mapped into the structural parameterization of the SVAR. By taking appropriate care of the volume elements, the draws were weighted so that they came from the normal-generalized-normal posterior distribution over the structural parameterization of the SVAR conditional on the identification restrictions. Since a proxy-SVAR identified with exogeneity restrictions can be represented by a SVAR identified with zero restrictions one may be tempted to use Arias, Rubio-Ramírez and Waggoner’s (2018) algorithm. However, the techniques of that paper cannot be directly applied in this environment because the zero restrictions embedded in a proxy-SVAR restrict its reduced-form representation, which invalidates the use of the orthogonal-reduced-form parameterization for our purpose. To handle this issue, we introduce a new parameterization called the orthogonal-triangular-block parameterization—composed of triangular-block parameters and orthogonal matrices—that makes it possible to sample independently from the structural parameterization of the proxy-SVAR even in the presence of the zero restrictions ingrained in this framework.

Using Waggoner and Zha (2003), we will produce independent draws from a normal-generalized-normal distribution over the triangular-block parameters and further generalize the QR decomposition to produce independent draws from a distribution over the orthogonal-triangular-block parameterization conditional on the exogeneity restrictions. We will then map these draws into the proxy-SVAR structural parameterization. These draws are weighted, by adapting the volume elements used in Arias, Rubio-Ramírez and Waggoner (2018), so that they come from the normal-generalized-normal posterior distribution over the structural parameterization of the proxy-SVAR conditional on the exogeneity restrictions of interest.

1For example, see Stock (2008), Stock and Watson (2012), Mertens and Ravn (2013), Gertler and Karadi (2015), and Montiel-Olea, Stock, and Watson (2016).

We also show that those restrictions may not be enough to identify the proxy-SVAR equations associated with structural shocks that are correlated with the proxies. In particular, additional sign and zero restrictions are needed for identification when more than one proxy is used to identify the same number of proxy-SVAR equations. For this reason, we adapt our main algorithm to consider these additional restrictions, which could be used not only to identify the proxy-SVAR equations associated with the structural shocks correlated with the proxies but also to identify the proxy-SVAR equations associated with those structural shocks that are uncorrelated with the proxies.

We present two applications to illustrate our algorithms. The first application is aimed at providing applied readers with a succinct and comprehensive description of how to use our techniques. To this end, we begin by studying the dynamic efects of consumption and investment total factor productivity (TFP) shocks in a proxy-SVAR where the equations associated with the shocks of interest are identified using Fernald’s (2014) TFP series as external instruments as in Lunsford (2016). An important diference between our approach and Lunsford’s (2016) approach is that, while he identifies one structural equation at a time by using a single instrument, we jointly identify two structural equations using two instruments.2 Hence, the application allows us to emphasize how to use additional sign and zero restrictions to simultaneously identify more than one equation. In particular, we identify the structural equations by assuming that they are the only equations whose shocks are correlated with the two external instruments and by adding some additional sign restrictions to parse out consumption TFP shocks from investment TFP shocks. Like Lunsford (2016), we find that a positive consumption TFP shock causes an increase in real GDP and consumption in non-durables and services as well as in durables and equipment while the price level gradually decreases. Accordingly, such a shock resembles a standard TFP shock. In contrast, a positive investment TFP shock leads to a decrease of real GDP, employment, consumption, and the price level. These results are inconsistent with the conventional wisdom of standard TFP shocks but in line with the findings in Liu, Fernald, and Basu (2012).

The second application is aimed at highlighting that the distinctive feature of our approach illustrated in the previous application—i.e., using more than one instrument to simultaneously set identify more than one structural equation—can provide critical insights for a few but highly influential studies identifying two structural equations using two instruments such as Mertens and Ravn (2013) and Mertens and Montiel-Olea (2018). As we will discuss later, the fundamental issue with their approach is that in order to separately identify the two structural equations of interest they are limited to consider a narrow class of additional zero restrictions that are hard to justify. We will make this clear by revisiting Mertens and Montiel-Olea (2018). This paper relies on proxy-SVARs to study the efects of exogenous changes in marginal and average personal income tax rates. One of their main conclusions is that, even though the response of reported income to exogenous changes in marginal tax rates is strong and significant, the response to exogenous changes in average tax rates is not statistically significant at any horizon. We will show that the identification scheme underlying this result exactly identifies the equations associated with both tax shocks by imposing an additional zero restriction on the systematic component of tax policies. We argue that their identification scheme is hard to justify; as a result, we substitute it for a set of less questionable sign restrictions. Once this is done, we find that both substitution and wealth efects play a relevant role for the transmission of tax rate shocks.

2Lunsford’s (2016) approach is a common approach in the literature (see Stock and Watson, 2012).

1.1 Relationship with the Bayesian Literature on Proxy-SVARs

As mentioned above, to the best of our knowledge, only three papers consider proxy-SVARs under the Bayesian paradigm: Bahaj (2014), Drautzburg (2016), and Caldara and Herbst (2016). Bahaj (2014) and Drautzburg (2016) draw from the posterior distribution of the orthogonal reduced-form parameterization and then transform the draws to the structural proxy-SVAR parameterization. Since the reduced-form parameters are restricted they both used Gibbs samplers to draw from the posterior of the reduced-form parameters; thereby the draws are not independent. But this is not the only diference with our approach. More importantly, they ignore the volume element when transforming the draws into the structural proxy-SVAR parameterization. As explained in Arias, Rubio-Ramírez, and Waggoner (2018), this implies that they are not drawing from a normal-generalized-normal posterior distribution over the structural parameterization of the proxy-SVAR conditional on the exogeneity restrictions and that the posterior distribution from which they are drawing depends on the zero restrictions—making comparisons across identification schemes impossible.

Finally, let’s relate our paper to Caldara and Herbst (2016). As in our approach, this paper draws from the structural parameterization of the proxy-SVAR. Nevertheless, while our approach imposes a conjugate prior on the structural parameterization, they work with non-conjugate prior densities, which complicates inference. In particular, the posterior draws are not independent and their Metropolis-Hastings sampler could become computationally ineficient compared with ours in large models. One advantage of Caldara and Herbst’s (2016)

approach relative to ours is that they can impose any desired prior on the strength of the proxy. Even so, researchers using our approach can also explore the sensitivity of the results to the quality of the instruments by imposing thresholds on their reliability using additional sign restrictions.

The remainder of the paper is organized as follows. Section 2 introduces the methodology. Section 3 describes the algorithm. Section 4 shows that when identifying more than one shock using more than one external instrument additional sign or zero restrictions are necessary, and presents an algorithm to consider them. Sections 5 and 6 present two applications. Section 7 concludes. Technical details are deferred to the Appendix.

2 The Framework

This section discusses our general framework. In Section 2.1, we describe the structural parameterization of a proxy-SVAR. In Section 2.2, we present the identification problem and the exogeneity restrictions. In Section 2.3, we explicitly specify the family of prior and posterior distributions over the structural parameterization of a proxy-SVAR that we would like to independently draw from. In Section 2.4, we introduce the orthogonal triangular-block parameterization. As will become clear later, this is a useful parameterization of a proxy-SVAR and our algorithms will rely on it. In this section we will also introduce the mapping that we will be using to move between parameterizations.

2.1 A Proxy-SVAR

Let be a vector of endogenous variables, be a vector of instruments (also called proxies), , and . If these are governed by a SVAR, then

\[\tilde {\boldsymbol {y}} _ {t} ^ {\prime} \tilde {\boldsymbol {A}} _ {0} = \sum_ {\ell = 1} ^ {p} \tilde {\boldsymbol {y}} _ {t - \ell} ^ {\prime} \tilde {\boldsymbol {A}} _ {\ell} + \tilde {\boldsymbol {c}} + \tilde {\varepsilon} _ {t} ^ {\prime} \mathrm{for} 1 \leq t \leq T,\tag{1}\]

where is an matrix for with invertible, c˜ is row vector, and is conditionally standard normal with mean zero and identity variance-covariance matrix.3 If and , Equation (1) can be more compactly written as

3We can always include a vector of exogenous variables, of dimension . In that case the model will be written as
p y˜0tA0 = X0t−`A + c˜ + z˜0td+ ε˜0t for 1 ≤ t ≤ T, `=1
where is row vector.

\[\tilde {\boldsymbol {y}} _ {t} ^ {\prime} \tilde {\boldsymbol {A}} _ {0} = \tilde {\boldsymbol {x}} _ {t} ^ {\prime} \tilde {\boldsymbol {A}} _ {+} + \tilde {\varepsilon} _ {t} ^ {\prime} \text { for } 1 \leq t \leq T.\tag{2}\]

Let , where is and is . Because is conditionally standard normal with mean zero and identity variance-covariance matrix, we are assuming that is uncorrelated with proxy-SVAR imposes that evolves according to

\[\pmb {y} _ {t} ^ {\prime} \pmb {A} _ {0} = \pmb {x} _ {t} ^ {\prime} \pmb {A} _ {+} + \pmb {\varepsilon} _ {t} ^ {\prime} \mathrm{for} 1 \leq t \leq T,\tag{3}\]

where and , with an matrix for invertible, and c a row vector.

Hence, are the structural shocks and are other shocks that afect the proxies. Equation (3) implies that

\[\tilde {\boldsymbol {A}} _ {i} = \left[ \begin{array}{c c} \boldsymbol {A} _ {i} & \boldsymbol {\Gamma} _ {i, 1} \\ \boldsymbol {0} _ {k \times n} & \boldsymbol {\Gamma} _ {i, 2} \end{array} \right],\]

where is and is for and is a matrix of zeros. We call the zero restrictions on and the block restrictions. We call Equation (2) plus the block restrictions the structural parameterization of the proxy-SVAR and such that the block restrictions hold the proxy-SVAR structural parameters. Hence, for the set of proxy-SVAR structural parameters the block restrictions always hold. Mertens and Ravn (2013) and Stock and Watson (2018) also model and as a joint system, but the structure is not exactly the same.

2.2 The Identification Problem in a Proxy-SVAR

Following Rothenberg (1971) the proxy-SVAR structural parameters and are observationally equivalent if and only if they imply the same joint distribution of . Proposition 1 extends an insight of Rubio-Ramírez, Waggoner, and Zha (2010) to the case of proxy-SVARs. Let denote a block diagonal matrix with the matrices along the diagonal.

Proposition 1. The proxy-SVAR structural parameters and are observationally equivalent and only if and , for some matrix , where is the set of

orthogonal matrices and Q is defined

\[\mathcal {Q} = \left\{\boldsymbol {Q} \in \mathcal {O} (\tilde {n}) | \boldsymbol {Q} = \operatorname{diag} \left(\boldsymbol {Q} _ {1}, \boldsymbol {Q} _ {2}\right), \boldsymbol {Q} _ {1} \in \mathcal {O} (n), a n d \boldsymbol {Q} _ {2} \in \mathcal {O} (k) \right\}.\]

Proof. See Appendix A.1.

Rubio-Ramírez, Waggoner, and Zha (2010) prove that SVARs are not identified. Corollary 1 shows that proxy-SVARs are not identified.

Corollary 1. A proxy-SVAR is not identified.

Proof. Because elements of are block diagonal, if are proxy-SVAR parameters, then so are for all . By Proposition 1, and are observationally equivalent. So proxy-SVARs are not identified. □

The identification problem in proxy-SVARs is commonly a partial identification problem because researchers focus on identifying a subset of the proxy-SVAR structural equations.4 For easy of exposition, henceforward we adopt Leeper, Sims, and Zha’s (1996) view and often talk about identifying structural shocks as equivalent to identifying structural equations because, under the proxy-SVAR framework, each equation contains only one shock so that it is possible to directly relate equations and shocks.5

More specifically, the identification problem in proxy-SVARs is typically solved by assuming that the k proxies are correlated with k structural shocks in and uncorrelated with the remaining structural shocks. Without loss of generality let the structural shocks correlated with the proxies be the last k elements of and the structural shocks uncorrelated with the proxies be the first elements of . We now show that the latter restrictions—which are known in the literature as exogeneity restrictions—are zero restrictions on a non-linear function of the proxy-SVAR structural parameters. To see this, first note that for the proxy-SVAR structural parameters it is the case that

\[\tilde {\boldsymbol {A}} _ {0} ^ {- 1} = \left[ \begin{array}{c c} \boldsymbol {A} _ {0} ^ {- 1} & - \boldsymbol {A} _ {0} ^ {- 1} \boldsymbol {\Gamma} _ {0, 1} \boldsymbol {\Gamma} _ {0, 2} ^ {- 1} \\ \boldsymbol {0} _ {k \times n} & \boldsymbol {\Gamma} _ {0, 2} ^ {- 1} \end{array} \right].\]

4A proxy-SVAR equation is identified if for any two sets of observationally equivalent proxy-SVAR parameters, the parameters in that equation are identical.
5As in Arias, Rubio-Ramírez and Waggoner (2018), the theory and the algorithms of this paper can be replicated for a proxy-SVAR in which the structural parameters are written in terms of impulse response functions (IRFs)—i.e., the IRF parameterization (see Appendix B in Arias, Rubio-Ramírez and Waggoner (2018)). In such a case, identifying structural shocks is equivalent to identifying structural IRFs.

Then, note that by multiplying Equation (2) by and focusing on the last k equations we obtain

\[\boldsymbol {m} _ {t} ^ {\prime} = \tilde {\boldsymbol {x}} _ {t} ^ {\prime} \tilde {\boldsymbol {A}} _ {+} \left[ \begin{array}{c} - \boldsymbol {A} _ {0} ^ {- 1} \boldsymbol {\Gamma} _ {0, 1} \boldsymbol {\Gamma} _ {0, 2} ^ {- 1} \\ \boldsymbol {\Gamma} _ {0, 2} ^ {- 1} \end{array} \right] - \varepsilon_ {t} ^ {\prime} \boldsymbol {A} _ {0} ^ {- 1} \boldsymbol {\Gamma} _ {0, 1} \boldsymbol {\Gamma} _ {0, 2} ^ {- 1} + \boldsymbol {\nu} _ {t} ^ {\prime} \boldsymbol {\Gamma} _ {0, 2} ^ {- 1} \text {for} 1 \leq t \leq T.\]

It follows that

\[\mathbb {E} [ \pmb {\varepsilon} _ {t} \pmb {m} _ {t} ^ {\prime} ] = - \pmb {A} _ {0} ^ {- 1} \pmb {\Gamma} _ {0, 1} \pmb {\Gamma} _ {0, 2} ^ {- 1}.\]

Thus, the identifying restrictions imply that the first rows of matrix must be zero. Taking transposes, this implies that the lower left-hand block of must be zero. This makes clear that proxy-SVARs are typically identified by zero restrictions on a function of the proxy-SVAR structural parameters. That is, the exogeneity restrictions are zero restrictions on . In addition to exogeneity restrictions, we also need the covariance matrix of the last k shocks and the k proxies, which is given by the last k rows of , to be non-singular. Following the literature, we refer to this as the relevance condition.

2.3 Priors, Posteriors, and a Useful Parameterization and Mapping

We will use a restricted normal-generalized-normal distribution over the structural parameterization of the proxy-SVAR as our prior distribution.6 Hence, the proxy-SVAR structural parameters have a prior density proportional to , where

\[N G N _ {(\nu , \Phi , \Psi , \Omega)} (\tilde {\boldsymbol {A}} _ {0}, \tilde {\boldsymbol {A}} _ {+}) \propto | \det (\tilde {\boldsymbol {A}} _ {0}) | ^ {\nu - n} e ^ {- \frac {1}{2} \operatorname{vec} (\tilde {\boldsymbol {A}} _ {0}) ^ {\prime} \Phi \operatorname{vec} (\tilde {\boldsymbol {A}} _ {0})} e ^ {- \frac {1}{2} (\operatorname{vec} (\tilde {\boldsymbol {A}} _ {+}) - \Psi \operatorname{vec} (\tilde {\boldsymbol {A}} _ {0})) ^ {\prime} \Omega^ {- 1} (\operatorname{vec} (\tilde {\boldsymbol {A}} _ {+}) - \Psi \operatorname{vec} (\tilde {\boldsymbol {A}} _ {0}))}.\tag{4}\]

The density is characterized by four parameters: a scalar , an symmetric and positive definite matrix Φ, an matrix Ψ, and an symmetric and positive definite matrix The normalgeneralized-normal distribution over the structural parameterization is a conjugate family of distributions commonly used in the literature. For instance, the Sims-Zha prior (see Sims and Zha, 1998) belongs to this family.

Our objective is to independently draw from the restricted normal-generalized-normal posterior distribution over the structural parametrization of a proxy-SVAR conditional on the exogeneity restrictions implied by Equation (4). More specifically, such posterior distribution is proportional to , where , and

6By a restricted normal-generalized-normal distribution over the structural parameterization of the proxy-SVAR we mean a normal-generalized-normal distribution over conditional on the block restrictions, where . If there are exogenous variables,
7Arias, Rubio-Ramírez, and Waggoner (2018) assumed a Kronecker structure for Φ, Ψ, and Ω.

Arias, Rubio-Ramírez, and Waggoner (2018) showed how to independently draw from a normal-generalizednormal posterior distribution over the structural parameterization of a SVAR conditional on sign and zero restrictions. Since a proxy-SVAR identified with exogeneity restrictions can be represented by the SVAR in Equation (2) identified with the zero restrictions on , and associated with the block and the exogeneity restrictions, one would like to use Arias, Rubio-Ramírez and Waggoner’s (2018) algorithm. However, the techniques of that paper cannot be directly applied in this context because the number of zero restrictions implied by the block restrictions alone is too large. There are block restrictions on each of the first n columns of , while the maximum number of restrictions that the aforementioned algorithm can handle on the column of the structural parameters is . So unless , an uninteresting case, the maximum will be exceeded for the column, if not before. In this paper we show that the techniques in Arias, Rubio-Ramírez, and Waggoner (2018) can be adapted to accomplish our objective.

The idea in Arias, Rubio-Ramírez, and Waggoner (2018) was to map independent draws from the orthogonalreduced-form parameterization conditional on the zero restrictions into the structural parameterization of the SVAR to create a proposal for the desired normal-generalized-normal posterior distribution over the structural parameterization conditional on the zero restrictions. The key to their approach is to properly account for the volume element associated with that mapping in order to characterize the proposal. This proposal was then embedded in an importance sampling algorithm.

The fact that the number of zeros is too large in a proxy-SVAR identified with exogeneity restrictions implies that the reduced-form is restricted. This prevents us from obtaining independent draws from the orthogonal reduced-form parameterization. Hence, instead of using the orthogonal-reduced-form parameterization, we will map independent draws from what we will call the orthogonal-triangular-block parameterization conditional on the exogeneity restrictions into the structural parameterization of the proxy-SVAR to create a proposal for the desired restricted normal-generalized-normal posterior distribution over the structural parameterization of the proxy-SVAR conditional on the exogeneity restrictions. As in Arias, Rubio-Ramírez, and Waggoner (2018), the key will be to properly account for the volume element in order to characterize the proposal. This proposal will be again used in an importance sampling algorithm.

2.4 The Orthogonal-Triangular-Block Parameterization and the Mapping

Let be an matrix, be an matrix, be an orthogonal matrix, and be a orthogonal matrix. The matrix is restricted to be upper-triangular with positive diagonal. The matrix , where is for and d is , is restricted so that the lower left-hand k ×n block of is zero for . We label the zero restrictions on and the triangular-block restrictions, and we call such that the triangular-block restrictions hold the triangular-block parameters.

Given any values of the triangular-block parameters and the orthogonal matrices , we can map into proxy-SVAR structural parameters ) by8

\[(\tilde {\boldsymbol {\Lambda}} _ {0}, \tilde {\boldsymbol {\Lambda}} _ {+}, \boldsymbol {Q} _ {1}, \boldsymbol {Q} _ {2}) \stackrel {{f}} {{\longrightarrow}} (\underbrace {\tilde {\boldsymbol {\Lambda}} _ {0} \operatorname{diag} (\boldsymbol {Q} _ {1} , \boldsymbol {Q} _ {2})} _ {\tilde {\boldsymbol {A}} _ {0}}, \underbrace {\tilde {\boldsymbol {\Lambda}} _ {+} \operatorname{diag} (\boldsymbol {Q} _ {1} , \boldsymbol {Q} _ {2})} _ {\tilde {\boldsymbol {A}} _ {+}}).\]

It is easy to verify that will satisfy the block restrictions , they are proxy-SVAR structural parameters).

The mapping has an inverse. Let be the QR-decomposition of normalized so that the diagonal of R is positive. Because the lower left-hand block of is zero, , where is and is . The inverse of is

\[(\tilde {A} _ {0}, \tilde {A} _ {+}) \stackrel {{f ^ {- 1}}} {{\longrightarrow}} (\underbrace {\tilde {A} _ {0} P} _ {\tilde {\Lambda} _ {0}}, \underbrace {\tilde {A} _ {+} P} _ {\tilde {\Lambda} _ {+}}, \underbrace {P _ {1} ^ {\prime}} _ {Q _ {1}}, \underbrace {P _ {2} ^ {\prime}} _ {Q _ {2}}).\]

The matrix will be upper triangular with positive diagonal because . Furthermore, since is block diagonal and the lower left-hand block of is , the lower left-hand block of each will be zero.

Just as a standard SVAR can alternatively be written in the orthogonal reduced-form parameterization, the triangular-block parameters together with the orthogonal matrices define an equivalent parameterization of the proxy-SVAR characterized by Equation (2) and the block restrictions. We call this alternative parameterization the orthogonal triangular-block parameterization of a proxy-SVAR and we write the latter as follows

\[\tilde {\pmb {y}} _ {t} ^ {\prime} \tilde {\pmb {\Lambda}} _ {0} = \tilde {\pmb {x}} _ {t} ^ {\prime} \tilde {\pmb {\Lambda}} _ {+} + \tilde {\pmb {u}} _ {t} ^ {\prime} \mathrm{for} 1 \leq t \leq T,\tag{5}\]

where with . Like , the innovations are conditionally standard normal.

8The function can be defined for all but will not be one-to-one over this larger set.

3 The Algorithm

In this section, we present an algorithm to make independent draws from the desired restricted normalgeneralized-normal posterior distribution over the structural parameterization of a proxy-SVAR conditional on the exogeneity restrictions. We achieve this goal by first independently drawing triangular-block parameters, , using Waggoner and Zha’s (2003) Gibbs sampler. Then, we show that the exogeneity restrictions are linear restrictions on the columns of the orthogonal matrix . This will allow us to use the ideas in Arias, Rubio-Ramírez and Waggoner (2018) to draw the orthogonal matrices , conditional on each draw of the triangular-block parameters, such that the exogeneity restrictions hold. Then, we use f to map triangularblock parameters plus the orthogonal matrices, , into proxy-SVAR structural parameters, . While these independent draws are not from the desired restricted normal-generalized-normal posterior distribution over the structural parameterization of a proxy-SVAR conditional on the exogeneity restrictions, we will be able to numerically compute the density associated with the implied distribution. Hence, we can use those draws as an intermediate step in an importance sampler to draw from the desired posterior distribution.

3.1 Independent Draws of the Triangular-Block Parameters

We use the Gibbs sampler of Waggoner and Zha (2003) to independently draw from a restricted normalgeneralized-normal posterior distribution over the triangular-block parameters characterized by .9 This Gibbs sampler can be used to draw from a normal-generalized-normal distribution subject to linear restrictions, as long as the restrictions do not involve cross-equation restrictions and the matrices , and are block diagonal.10 Since the triangular and block restrictions on are exclusion restrictions that do not involve cross-equation restrictions, and , and can be chosen to be block diagonal, the conditions for using the Gibbs sampler are satisfied. Furthermore, because is restricted to be upper-triangular, it follows from Theorem 2 of Waggoner and Zha (2003) that the Gibbs sampler draws will be independent. In Appendix A.4, we describe how to adapt their paper to our purposes.

Often, it sufices to choose to be equal to , the parameters associated with the desired restricted normal-generalized-normal posterior distribution over the structural parameterization of the proxy-SVAR conditional on the exogeneity restrictions. However, sometimes this can lead to small efective sample sizes in our importance sampler. In Appendix A.5, we describe a more tailored choice of that can avoid this loss of eficiency.

10The Gibbs sampler of Waggoner and Zha (2003) was developed to draw from the posterior distribution of a structural VAR with linear non-cross-equation restrictions using a certain class of normal priors. The class of posterior distributions that can be obtained with this class of priors is the set of all normal-generalized-normal distributions with , Φ block diagonal with n˜ symmetric and positive definite blocks, block diagonal with n˜ arbitrary blocks, and block diagonal with n˜ symmetric and positive definite m˜ × m˜ blocks, conditional on the linear non-cross-equation restrictions. In Appendix A.4 we outline this Gibbs sampler in our context.
9Here, by a restricted normal-generalized-normal distribution over the triangular-block parameters we mean a normalgeneralized-normal distribution over conditional on the triangular-block restrictions.

3.2 Exogeneity Restrictions in the Orthogonal-Triangular-Block Parameterization

Let and . Because of the arguments made in Section 2.2, if are proxy-SVAR structural parameters, the exogeneity restrictions are of the form

\[\boldsymbol {J} (\tilde {\boldsymbol {A}} _ {0} ^ {- 1}) ^ {\prime} \boldsymbol {e} _ {\tilde {n}, j} = \mathbf {0} _ {k \times 1} \text { for } 1 \leq j \leq n - k.\tag{6}\]

The index stops at because there are no exogeneity restrictions for . In terms of the orthogonal-triangular-block parameterization, this is equivalent to

\[\boldsymbol {J} (\tilde {\boldsymbol {\Lambda}} _ {0} ^ {- 1}) ^ {\prime} \operatorname{diag} (\boldsymbol {Q} _ {1}, \boldsymbol {Q} _ {2}) \boldsymbol {e} _ {\tilde {n}, j} = \underbrace {\boldsymbol {J} (\tilde {\boldsymbol {\Lambda}} _ {0} ^ {- 1}) ^ {\prime} \boldsymbol {L} ^ {\prime}} _ {\boldsymbol {G} (\tilde {\boldsymbol {\Lambda}} _ {0})} \boldsymbol {Q} _ {1} \boldsymbol {e} _ {n, j} = \mathbf {0} _ {k \times 1} \text {for} 1 \leq j \leq n - k,\tag{7}\]

where we have used the fact that . Thus, conditional on a draw of triangular-block parameters , the exogeneity restrictions are linear restrictions on the columns of . We will denote the number of exogeneity restrictions on the column of by , which is k if and is zero if . As in Arias, Rubio-Ramírez, and Waggoner (2018), we will use this fact to draw the orthogonal matrices . As will become clear below, drawing the orthogonal matrix is simpler because its columns are unrelated to the exogeneity restrictions.

3.3 An Algorithm

Sections 3.1 and 3.2 suggest that we can devise an algorithm to make independent draws from a distribution over the structural parameterization of a proxy-SVAR conditional on the exogeneity restrictions.

Algorithm 1. The following algorithm makes independent draws from a distribution over the structural parameterization of a proxy-SVAR conditional on the exogeneity restrictions.

1. Draw triangular-block parameters independently from the restricted distribution using Waggoner and Zha’s (2003) Gibbs sampler.

2. For , draw independently from a standard normal distribution and set

3. For , draw independently from a standard normal distribution and set

4. Define recursively by for any matrix whose columns form an orthonormal basis for the null space of the matrix:

\[M _ {1, j} = \left\{ \begin{array}{l l} \left[ \begin{array}{c c c c} \boldsymbol {q} _ {1, 1} & \dots & \boldsymbol {q} _ {1, j - 1} & \boldsymbol {G} (\tilde {\boldsymbol {\Lambda}} _ {0}) ^ {\prime} \end{array} \right] ^ {\prime} & f o r 1 \leq j \leq n - k \\ \left[ \begin{array}{c c c} \boldsymbol {q} _ {1, 1} & \dots & \boldsymbol {q} _ {1, j - 1} \end{array} \right] ^ {\prime} & f o r n - k + 1 \leq j \leq n. \end{array} \right.\]

5. Define recursively by for any matrix whose columns form an orthonormal basis for the null space of the matrix:

\[\boldsymbol {M} _ {2, j} = \left[ \begin{array}{c c c} \boldsymbol {q} _ {2, 1} & \dots & \boldsymbol {q} _ {2, j - 1} \end{array} \right] ^ {\prime} f o r 1 \leq j \leq k.\]

6. Set

7. Return to Step 1 until the required number of draws has been obtained.

In order for this algorithm to work, it must be the case that is of full row rank; otherwise the number of columns in will not equal the dimension of and so the product will not be defined. When and or when and , the matrix will clearly be of full row rank. However, when and , the matrix will be of full row rank if and only if is of full row rank. This is because, by construction, the are perpendicular to . Since the probability of drawing a such that the is not of full row rank is zero, we can assume without loss of generality that is of full row rank.

When the exogeneity restrictions hold, the relevance condition is equivalent to being of full row rank. To see this, note that , so that by properties of the rank is of full row rank if and only if is of full row rank. If the exogeneity restrictions hold, then

\[\mathbb {E} [ \boldsymbol {m} _ {t} \boldsymbol {\varepsilon} _ {t} ^ {\prime} ] = - \left(\boldsymbol {A} _ {0} ^ {- 1} \boldsymbol {\Gamma} _ {0, 1} \boldsymbol {\Gamma} _ {0, 2} ^ {- 1}\right) ^ {\prime} = \boldsymbol {J} (\tilde {\boldsymbol {A}} _ {0} ^ {- 1}) ^ {\prime} \boldsymbol {L} ^ {\prime} = [ \boldsymbol {0} _ {k \times (n - k)} \boldsymbol {V} ],\]

where the matrix V is the covariance matrix of the k proxy variables and the last k structural shocks. So, the relevance condition, which requires V to be non-singular, holds if and only if is of full row rank.

Of course, in practice, we not only want the covariance matrix to be non-singular, but we also would like it to be well conditioned so that it is far from being singular.

It is also the case that the matrix is not unique. If the columns of form an orthonormal basis for the null space of , then so will the columns of for any orthogonal matrix X. The particular choice of does not make a material diference in the output of Algorithm 1, but in the next section, when we compute the density over the structural parameterization of a proxy-SVAR conditional on the exogeneity restrictions implied by Algorithm 1, we will need to be able to define the function so that it is continuously diferentiable over a suitable large set. In Appendix A.2, we describe a particular choice that will work.

The independent draws of the proxy-SVAR structural parameters conditional on the exogeneity restrictions produced by Algorithm 1 will not be from the desired restricted normal-generalized-normal posterior distribution over the structural parameterization of a proxy-SVAR conditional on the exogeneity restrictions. The density implied by Algorithm 1 will be analyzed in Section 3.4 to follow. It is important to clarify that our algorithm can handle cases in which a researcher wants to consider k instruments that are correlated with shocks, with . In such cases, Equations (6) and (7) will only hold for . This could be of interest for example when a researcher assumes that a proxy is not correlated with a particular structural shock while leaving the correlation with the remaining shocks unrestricted.

3.4 The Density Implied by Algorithm 1

Step 1 of Algorithm 1 independently draws triangular-block parameters from a restricted normal generalized-normal. Step 2 draws from the uniform distribution on the unit sphere in Step 3 draws from the uniform distribution on the unit sphere in . Hence, the density over is proportional to a restricted normal-generalized-normal. Steps 4 and 5 map , we denote this mapping by g. Finally, Step 6 maps into proxy-SVAR structural parameters using the function f. The composite mapping implied by Steps 1 through Step 6 together with Theorem 3 in Arias, Rubio-Ramírez, and Waggoner (2018) will be used to compute the density implied by Algorithm 1.

Since as noted before the matrix is not unique, the function g is not uniquely defined. In Appendix A.2 we will show that g can be defined so that it is one-to-one and continuously diferentiable except on a closed set of measure zero. This is what is needed to apply Theorem 3 in Arias, Rubio-Ramírez, and Waggoner (2018).

If Z denotes the set of all proxy-SVAR structural parameters that satisfy the exogeneity restrictions and the relevance condition, then by Theorem 3 in Arias, Rubio-Ramírez, and Waggoner (2018), the density over the structural parameterization of the proxy-SVAR conditional on the exogeneity restrictions implied by Algorithm 1 is proportional to

\[N G N _ {(\hat {\nu}, \hat {\Phi}, \hat {\Psi}, \hat {\Omega})} (\tilde {\Lambda} _ {0}, \tilde {\Lambda} _ {+}) v _ {(f \circ g) ^ {- 1} | _ {\mathcal {Z}}} (\tilde {A} _ {0}, \tilde {A} _ {+}),\]

where

3.5 An Importance Sampler

The results above show that Algorithm 1 generates independent draws from a distribution over the proxy-SVAR structural parameterization conditional on the exogeneity restrictions that is not equal to the desired restricted normal-generalized-normal posterior distribution over the proxy-SVAR structural parameterization conditional on the exogeneity restrictions. However, because Section 3.4 shows how to numerically evaluate the former distribution, we can use such a distribution as a proposal in the following importance sampler algorithm to accomplish our objective.

Algorithm 2. The following algorithm independently draws from the restricted posterior distribution over the proxy-SVAR structural parameterization conditional on the exogeneity restrictions and the relevance condition.

1. Use Algorithm 1 to independently draw proxy-SVAR structural parameters that satisfy the exogeneity restrictions and the relevance condition.

2. Set its importance weight to

\[\frac {N G N _ {(\tilde {\nu} , \tilde {\Phi} , \tilde {\Psi} , \tilde {\Omega})} (\tilde {\boldsymbol {A}} _ {0} , \tilde {\boldsymbol {A}} _ {+})}{N G N _ {(\hat {\nu} , \hat {\Phi} , \hat {\Psi} , \hat {\Omega})} (\tilde {\boldsymbol {\Lambda}} _ {0} , \tilde {\boldsymbol {\Lambda}} _ {+}) v _ {(f \circ g) ^ {- 1} | _ {\mathcal {Z}}} (\tilde {\boldsymbol {A}} _ {0} , \tilde {\boldsymbol {A}} _ {+})},\]

where and denotes the set of all proxy-SVAR structural parameters that satisfy the exogeneity restrictions and the relevance condition.

3. Return to Step 1 until the required number of draws has been obtained.

4. Re-sample with replacement using the importance weights.

Step 2 is the crucial one. The re-sampling Step 4 allows us to have unweighted and independent draws. Given the desired number of independent draws, the researcher should require enough draws from Steps 1-3 so that the efective sample size is at least as large as the number of desired independent draws. We define the

efective sample size as

\[N \left(\sum_ {i = 1} ^ {N} w _ {i}\right) ^ {2} / \left(\sum_ {i = 1} ^ {N} w _ {i} ^ {2}\right),\]

where is the weight associated with the draw and N is the total number of draws obtained in Steps 1-3.

Algorithm 2 is stated in terms of the proxy-SVAR structural parameterization, but it will work for any parameterization as long as one can explicitly compute the transformation between such parameterization and the orthogonal triangular-block parameterization. It is also important to note that computing the volume element in Step 2 is the most expensive part in implementing Algorithm 2. The rest of Algorithm 2 is quite fast.11

4 The Need for Additional Restrictions

We have described an algorithm that allows us to independently draw from the desired restricted normalgeneralized-normal posterior distribution over the proxy-SVAR structural parameterization conditional on the exogeneity restrictions. Next, we show that the exogeneity restrictions only allow us to categorize the proxy-SVAR shocks into two groups: the ones that are correlated with the proxies and the ones that are not correlated with the proxies. If we only use the exogeneity restrictions, we have an identification problem of the proxy-SVAR shocks that are correlated with the proxies; unless . The same problem occurs within the proxy-SVAR shocks that are not correlated with the proxies.

Proposition 2. Let and be proxy-SVAR structural parameters that also satisfy the exogeneity restrictions and the relevance condition, then and are observationally equivalent if and only if there exists a matrix such that and , where X is defined by

\[\mathcal {X} = \left\{\boldsymbol {Q} \in \mathcal {Q} | \boldsymbol {Q} = \operatorname{diag} \left(\boldsymbol {Q} _ {3}, \boldsymbol {Q} _ {4}, \boldsymbol {Q} _ {5}\right), \boldsymbol {Q} _ {3} \in \mathcal {O} (n - k), \boldsymbol {Q} _ {4} \in \mathcal {O} (k), a n d \boldsymbol {Q} _ {5} \in \mathcal {O} (k) \right\}.\]

Proof. See Appendix A.3.

Proposition 2 tells us that we need additional identification restrictions to identify the structural shocks within the set of structural shocks that are correlated with the proxies. In addition, the proposition makes clear that if we want to identify structural shocks within the set of the structural shocks that are not correlated with the proxies, then we also need additional identification restrictions. This is easy to see because will rotate the columns of the proxy-SVAR structural parameters associated with the structural shocks that are not correlated with the proxies while will rotate the proxy-SVAR equations associated with the structural shocks that are correlated with the proxies.12 The additional restrictions can be sign and/or zero restrictions. Since imposing additional sign restrictions is straightforward, we will first focus on additional zero restrictions by showing how to modify Algorithm 1 to incorporate them. Then, we will show how to modify Algorithm 2 to consider both additional sign and zero restrictions.

11The reader should note that Algorithm 2 must be used with a normalization. Typically, SVARs are normalized by restricting the sign of the response of a given variable to a shock of interest. Proxy-SVARs can be normalized analogously. Since this is well understood and one simply has to set the importance weight to zero when the normalization is not satisfied, we do not explicitly state this in the algorithm.

Often, one is interested only in partial identification of the k structural shocks that are correlated with the k proxies. If that is the case and , the exogeneity restrictions exactly identify the structural shock correlated with the proxy. The following corollary of Proposition 2 formalizes this result.

Corollary 2. Let and let be proxy-SVAR structural parameters that also satisfy the exogeneity restrictions and the relevance condition. Then, the last column of is identified up to a sign.

Proof. If , then we have and the column of Q is equal to

Notwithstanding, the theory and techniques that we develop apply to additional sign and zero restrictions on any function from the set of proxy-SVAR structural parameters to the set of matrices that satisfies the condition , for every 13

4.1 Additional Zero Restrictions

To set the notation, let be a matrix of full row rank, where for and for . In particular, we assume that the additional zero restrictions in the proxy-SVAR structural parameterization are of the form for . The fact that for reflects that we have already imposed the exogeneity restrictions on the first structural shocks. Note that we are identifying only . In principle, one could also use the additional zero restrictions to identify , but we do not believe that is of interest.

From the definition of f and the fact that

\[\boldsymbol {F} \left(\tilde {\boldsymbol {A}} _ {0} \operatorname{diag} \left(\boldsymbol {Q} _ {1}, \boldsymbol {Q} _ {2}\right), \tilde {\boldsymbol {A}} _ {+} \operatorname{diag} \left(\boldsymbol {Q} _ {1}, \boldsymbol {Q} _ {2}\right)\right) = \boldsymbol {F} \left(\tilde {\boldsymbol {A}} _ {0}, \tilde {\boldsymbol {A}} _ {+}\right) \operatorname{diag} \left(\boldsymbol {Q} _ {1}, \boldsymbol {Q} _ {2}\right),\]

13In addition, a regularity condition on is needed. For instance, it sufices to assume that is diferentiable and that its derivative is of full row rank.
12Note that and are related to introduced in Proposition 1 while is related to of the same proposition.

the zero restrictions in the orthogonal-triangular-block parameterization are

\[\boldsymbol {Z} _ {j} \boldsymbol {F} (f (\tilde {\boldsymbol {\Lambda}} _ {0}, \tilde {\boldsymbol {\Lambda}} _ {+}, \boldsymbol {Q} _ {1}, \boldsymbol {Q} _ {2})) \boldsymbol {e} _ {\tilde {n}, j} = \boldsymbol {Z} _ {j} \boldsymbol {F} (f (\tilde {\boldsymbol {\Lambda}} _ {0}, \tilde {\boldsymbol {\Lambda}} _ {+}, \boldsymbol {I} _ {n}, \boldsymbol {I} _ {k})) \operatorname{diag} (\boldsymbol {Q} _ {1}, \boldsymbol {Q} _ {2}) \boldsymbol {e} _ {\tilde {n}, j} = \mathbf {0} _ {z _ {j} \times 1} \text {for} 1 \leq j \leq n.\]

This last equation can be written as

\[\underbrace {\boldsymbol {Z} _ {j} \boldsymbol {F} (f (\tilde {\boldsymbol {\Lambda}} _ {0} , \tilde {\boldsymbol {\Lambda}} _ {+} , \boldsymbol {I} _ {n} , \boldsymbol {I} _ {k})) \boldsymbol {L} ^ {\prime}} _ {\boldsymbol {G} _ {1, j} (\tilde {\boldsymbol {\Lambda}} _ {0}, \tilde {\boldsymbol {\Lambda}} _ {+})} \boldsymbol {Q} _ {1} \boldsymbol {e} _ {n, j} = \boldsymbol {0} _ {z _ {j} \times 1} \text {for} 1 \leq j \leq n.\]

This allows us to modify Algorithm 1 to consider the additional zero restrictions as follows.

Algorithm 3. The following algorithm makes independent draws from a distribution over the structural parameterization of a proxy-SVAR conditional on the exogeneity restrictions and the additional zero restrictions.

1. Draw triangular-block parameters independently from the restricted distribution using Waggoner and Zha’s (2003) Gibbs sampler.

2. For , draw independently from a standard normal distribution and set

3. For , draw independently from a standard normal distribution and set

4. Define recursively by for any matrix whose columns form an orthonormal basis for the null space of the matrix:

\[M _ {1, j} = \left\{ \begin{array}{l l} \left[ \begin{array}{c c c c} \boldsymbol {q} _ {1, 1} & \dots & \boldsymbol {q} _ {1, j - 1} & \boldsymbol {G} (\tilde {\boldsymbol {\Lambda}} _ {0}) ^ {\prime} \\ & & & \boldsymbol {G} _ {1, j} (\tilde {\boldsymbol {\Lambda}} _ {0}, \tilde {\boldsymbol {\Lambda}} _ {+}) ^ {\prime} \end{array} \right] ^ {\prime} & f o r 1 \leq j \leq n - k \\ \left[ \begin{array}{c c c c} \boldsymbol {q} _ {1, 1} & \dots & \boldsymbol {q} _ {1, j - 1} & \boldsymbol {G} _ {1, j} (\tilde {\boldsymbol {\Lambda}} _ {0}, \tilde {\boldsymbol {\Lambda}} _ {+}) ^ {\prime} \end{array} \right] ^ {\prime} & f o r n - k + 1 \leq j \leq n. \end{array} \right.\]

5. Define recursively by for any matrix whose columns form an orthonormal basis for the null space of the matrix:

\[\boldsymbol {M} _ {2, j} = \left[ \begin{array}{c c c} \boldsymbol {q} _ {2, 1} & \dots & \boldsymbol {q} _ {2, j - 1} \end{array} \right] ^ {\prime} f o r 1 \leq j \leq k.\]

6. Set

7. Return to Step 1 until the required number of draws has been obtained.

As was the case with Algorithm 1, the independent draws of the proxy-SVAR structural parameters produced by Algorithm 3 will not be from the desired restricted normal-generalized-normal posterior distribution over the structural parameterization of a proxy-SVAR conditional on the exogeneity restrictions and the additiona zero restrictions. If X denotes the set of all proxy-SVAR structural parameters that satisfy the exogeneity restrictions, the relevance condition, and the additional zero restrictions, because of the results in Section 3.4, the density over the structural parameterization implied by Algorithm 3 is proportional to

\[N G N _ {(\hat {\nu}, \hat {\Phi}, \hat {\Psi}, \hat {\Omega})} (\tilde {\Lambda} _ {0}, \tilde {\Lambda} _ {+}) v _ {(f \circ g) ^ {- 1} | \mathcal {X}} (\tilde {A} _ {0}, \tilde {A} _ {+}),\]

where and has been adapted to consider the new matrices and Thus, we can use this density as our new proposal density. Thus, we next show how to adapt the importance sampler described in Algorithm 2 to accomplish our objective of making independent draws from the desired restricted normal-generalized-normal posterior distribution conditional on the exogeneity restrictions, the relevance condition, and the additional sign and zero restrictions.

To conclude this section, let’s clarify that our methodology cannot be used to impose the restriction that every proxy is only correlated with a particular structural shock. This requires imposing a diagonal structure in the matrix V using additional zero restrictions, which is outside the scope of our algorithms given the constraints embedded in for

4.2 Additional Sign and Zero Restrictions

We assume that the additional sign restrictions in the proxy-SVAR structural parameterization are of the form , where is continuous and is the number of sign restrictions. As explained in Arias, Rubio-Ramírez, and Waggoner (2018), because is continuous, the set of all proxy-SVAR structural parameters satisfying the exogeneity restrictions and the additional sign and zero restrictions will be open in the set of all proxy-SVAR structural parameters satisfying the exogeneity restrictions and the additional zero restrictions. Thus, if the additional sign restrictions are non-degenerate, so that there is at least one value of the proxy-SVAR structural parameters satisfying the exogeneity restrictions and the additional sign and zero restrictions, then the set of all proxy-SVAR structural parameters satisfying the exogeneity restrictions and the additional sign and zero restrictions will be of positive volume measure in the set of all proxy-SVAR structural parameters satisfying the exogeneity restrictions and the additional zero restrictions. This justifies algorithms of the type described below to accomplish our objective of making independent draws from the desired restricted normal-generalized-normal posterior distribution conditional on the exogeneity restrictions, the relevance condition, and the additional sign and zero restrictions. The next algorithm is a modification of

14The volume measure is defined in Arias, Rubio-Ramírez, and Waggoner (2018).

Algorithm 2 to consider additional sign and zero restrictions when using Algorithm 3 instead of Algorithm 1 in Step 1.

Algorithm 4. The following algorithm independently draws from the desired restricted posterior distribution over the proxy-SVAR structural parameterization conditional on the exogeneity restrictions, the relevance condition, and the additional sign and zero restrictions.

1. Use Algorithm 3 to independently draw proxy-SVAR structural parameters that satisfy the exclusion restrictions, the relevance condition, and the additional zero restrictions.

2. satisfies the sign restrictions, set its importance weight to

\[\frac {N G N _ {(\tilde {\nu} , \tilde {\Phi} , \tilde {\Psi} , \tilde {\Omega})} (\tilde {\boldsymbol {A}} _ {0} , \tilde {\boldsymbol {A}} _ {+})}{N G N _ {(\hat {\nu} , \hat {\Phi} , \hat {\Psi} , \hat {\Omega})} (\tilde {\boldsymbol {\Lambda}} _ {0} , \tilde {\boldsymbol {\Lambda}} _ {+}) v _ {(f \circ g) ^ {- 1} | \mathcal {X}} (\tilde {\boldsymbol {A}} _ {0} , \tilde {\boldsymbol {A}} _ {+})},\]

where and where X denotes the set of all proxy-SVAR structural parameters that satisfy the exclusion restrictions, the relevance condition, and the additional zero restrictions.

3. Return to Step 1 until the required number of draws has been obtained.

4. Re-sample with replacement using the importance weights.

As was the case with Algorithm 2, computing the volume element in Step 2 is the most expensive part in implementing Algorithm 4. The rest of Algorithm 4 is quite fast. But the reader should note that we do not need to compute the volume element for all the draws, only for those that satisfy the sign restrictions. Finally, as before we also need the relevance condition to hold.

5 Application I: The Dynamic Efects of TFP Shocks

In this section we illustrate our methodology by studying the dynamic efects of two types of TFP shocks, a consumption TFP shock and an investment TFP shock, in a quarterly frequency proxy-SVAR featuring five endogenous variables and two proxies for the shocks of interest. More specifically, we adopt the specification of the SVAR and the proxies from Lunsford (2016). Accordingly, the endogenous variables are real GDP growth, employment growth, inflation, real consumption growth, and real investment in equipment growth. The remaining details on the data are provided in Appendix A.6. The proxies are a consumption TFP proxy and an investment TFP proxy based on Fernald’s (2014) consumption and investment TFP series, respectively. In particular, we use Lunsford’s (2016) proxies, which are obtained by regressing each of the TFP series just mentioned on four lags of the endogenous variables and by labeling the residuals associated with each of these regressions as consumption and investment TFP proxies, respectively.15

The proxy-SVAR features four lags and a constant, and the sample runs from 1947Q2 until 2015Q4. Consequently, in this application , and . We set , and to characterize our prior over the proxy-SVAR structural parameters, and we set νˆ = ν˜, Φ = Φ , Φ = Φ and to characterize our proposal over the orthogonal-triangular-block parameterization.

Figura
Figura
Figura
Figura

Figure 1: IRFs to a positive one standard deviation consumption and investment TFP shocks. The blue solid-dotted curves represent the point-wise posterior medians and the gray shaded areas represent the 68 percent equal-tailed point-wise probability bands to a consumption TFP shock. The red solid curves represent the point-wise posterior medians and the red shaded areas represent the 68 percent equal-tailed point-wise probability bands to an investment TFP shock. The figure is based on 10,000 independent efective draws obtained using Algorithm 4.

Figure 1: IRFs to a positive one standard deviation consumption and investment TFP shocks. The blue solid-dotted curves represent the point-wise posterior medians and the gray shaded areas represent the 68 percent equal-tailed point-wise probability bands to a consumption TFP shock. The red solid curves represent the point-wise posterior medians and the red shaded areas represent the 68 percent equal-tailed point-wise probability bands to an investment TFP shock. The figure is based on 10,000 independent efective draws obtained using Algorithm 4.

Let be a vector containing the consumption and investment TFP shocks, i.e. and let be a vector containing all other structural shocks. The first set of identification assumptions are

\[\mathbb {E} \left[ \pmb {m} _ {t} \pmb {\varepsilon} _ {t} ^ {T F P \prime} \right] = \pmb {V} \neq \pmb {0} _ {2 \times 2} \mathrm{and} \mathbb {E} \left[ \pmb {m} _ {t} \pmb {\varepsilon} _ {t} ^ {O \prime} \right] = \pmb {0} _ {2 \times 3},\]

where are the proxies for the consumption and investment TFP shocks. As mentioned in

15We downloaded the proxies from Kurt Lunsford’s website at https://sites.google.com/site/kurtglunsford/research.

Section 4, without additional restrictions these conditions are not enough to distinguish a consumption TFP shock from an investment TFP shock. As a consequence we also impose the sign restrictions 2 and on the entries of V . If we order the two structural shocks of interest last, this implies setting

\[\boldsymbol {S} _ {4} = \left[ \begin{array}{c c c c c c c} 0 & 0 & 0 & 0 & 0 & 1 & 0 \\ 0 & 0 & 0 & 0 & 0 & 0 & 1 \end{array} \right], \boldsymbol {S} _ {5} = \left[ \begin{array}{c c c c c c c} 0 & 0 & 0 & 0 & 0 & 0 & 1 \\ 0 & 0 & 0 & 0 & 0 & 1 & 0 \end{array} \right],\]

and

\[\tilde {\boldsymbol {F}} (\tilde {\boldsymbol {A}} _ {0}, \tilde {\boldsymbol {A}} _ {+}) = \left[ \begin{array}{c} \boldsymbol {e} _ {2, 1} ^ {\prime} \boldsymbol {S} _ {4} (\tilde {\boldsymbol {A}} _ {0} ^ {- 1}) ^ {\prime} \boldsymbol {e} _ {\tilde {n}, 4} \\ \boldsymbol {e} _ {2, 1} ^ {\prime} \boldsymbol {S} _ {4} (\tilde {\boldsymbol {A}} _ {0} ^ {- 1}) ^ {\prime} \boldsymbol {e} _ {\tilde {n}, 4} - \boldsymbol {e} _ {2, 2} ^ {\prime} \boldsymbol {S} _ {5} (\tilde {\boldsymbol {A}} _ {0} ^ {- 1}) ^ {\prime} \boldsymbol {e} _ {\tilde {n}, 5} \\ \boldsymbol {e} _ {2, 1} ^ {\prime} \boldsymbol {S} _ {5} (\tilde {\boldsymbol {A}} _ {0} ^ {- 1}) ^ {\prime} \boldsymbol {e} _ {\tilde {n}, 5} \\ \boldsymbol {e} _ {2, 1} ^ {\prime} \boldsymbol {S} _ {5} (\tilde {\boldsymbol {A}} _ {0} ^ {- 1}) & \boldsymbol {e} _ {\tilde {n}, 5} ^ {\prime} - \boldsymbol {e} _ {2, 2} ^ {\prime} \boldsymbol {S} _ {4} (\tilde {\boldsymbol {A}} _ {0} ^ {- 1}) ^ {\prime} \boldsymbol {e} _ {\tilde {n}, 4} \end{array} \right].\]

Since in this application there are sign and exogeneity restrictions involved, we use Algorithm 4. More specifically, we obtain an efective sample size of 10,000 draws of the proxy-SVAR structural parameters satisfying the sign and exogeneity restrictions. Figure 1 shows the IRFs to a positive one standard deviation consumption and investment TFP shocks, respectively.16 The blue solid-dotted curves represent the point-wise posterior medians and the gray shaded areas represent the 68 percent equal-tailed point-wise probability bands to a consumption TFP shock. The red solid curves represent the point-wise posterior medians and the red shaded areas represent the 68 percent equal-tailed point-wise probability bands to an investment TFP shock. These IRFs are qualitatively consistent with the results reported by Lunsford (2016).

In particular, a consumption TFP shock causes an increase in real GDP, consumption in non-durables and services, consumption in durables and equipment, and employment while the price level gradually decreases. Although the probability bands associated with the latter variable contain zero, the findings are in line with those reported by Lunsford (2016). Accordingly, a consumption TFP shock implies opposite movements in quantities and prices supporting the conventional wisdom about the efects of standard TFP shocks. In contrast, a positive investment TFP shock leads with high probability to a decrease in real GDP, employment, consumption, and the price level. As highlighted by Lunsford (2016), these results are inconsistent with the conventional wisdom of standard TFP shocks but in line with the findings in Liu, Fernald, and Basu (2012).

Finally, the reliable use of the importance sampler requires the importance weights to possess finite variance. We use the tests proposed by Koopman, Shephard, and Creal (2009) as described in Appendix A.7. In

16While Lunsford (2016) reports the IRFs of the endogenous variables in the SVAR, we report the cumulative IRFs for easy of exposition.

Appendix A.8 we show that these tests imply that the finite variance requirement holds for the application analyzed in this section.

6 Application II: The Dynamic Efects of Personal Income Tax Shocks

In this section we use our methodology to revisit a recent study by Mertens and Montiel-Olea (2018) presenting new time series evidence—based on proxy-SVARs—on the efects of personal income tax rate cuts on reported income and other indicators of real activity such as GDP and the unemployment rate. The study under analysis works in the frequentist paradigm and its three main reported findings can be summarized as follows. First, negative average marginal tax rate (AMTR) shocks lead not only to increases in real GDP and declines in the unemployment rate but also to increases in reported income. Second, while income responds to AMTR shocks, it does not react to average tax rate (ATR) shocks, indicating that substitution efects instead of wealth efects are crucial for the transmission of tax rate shocks. Third, negative AMTR shocks for taxpayers at the top 1 percent of the income distribution have positive short-run efects on income and economic activity but zero long-run efects (i.e., 4 to 5 years after the shock). In contrast, negative AMTR shocks for tax payers at the bottom 99 percent of the income distribution have zero short-run efects on income and economic activity but positive long-run efects.

Our Bayesian approach will basically replicate Mertens and Montiel-Olea’s (2018) first finding. When analyzing their second finding, we will show that the identification scheme in Mertens and Montiel-Olea (2018) imposes a zero restriction on . More specifically, they exactly identify both the AMTR and ATR shocks by imposing a zero restriction on the systematic component of tax policies. It is important to note that their zero restriction is linked to a particular Cholesky factorization and, hence, the order of the variables afects the identification. Each ordering implies a diferent zero restriction on the systematic component of tax policy and, therefore, a diferent identification of the AMTR and ATR shocks. So, Mertens and Montiel-Olea (2018) are in fact using two diferent identification schemes. Strangely, when analyzing the IRFs for the AMTR shocks Mertens and Montiel-Olea (2018) use a diferent identification than when analyzing the IRFs for the ATR shocks. In any case, we think that both of their identification schemes are hard to justify because of the zero restriction on the systematic component of tax policies. Our methods will allow us to substitute this zero restriction for a set of less restrictive sign restrictions to be described below. Accordingly, we set identify the AMTR and ATR shocks. Once this is done, and contrary to what Mertens and Montiel-Olea (2018) report, we find that both substitution and wealth efects seem to play a relevant role for the transmission of tax rate shocks. Because the IRFs for AMTR shocks implied by both orderings (identification schemes) are in agreement, one could be tempted to conclude that Mertens and Montiel-Olea’s (2018) results are robust; in fact, Mertens and Montiel-Olea (2018) seem to fall into this trap by comparing Panels (A) and (B) in Figure 10 of their paper and highlighting that the results are similar. An additional insight of our exercise is that just checking two possible identification schemes is a dangerous practice that can give the researcher a false sense of robustness on her results.

When analyzing Mertens and Montiel-Olea’s (2018) third finding, we first show that they place a zero restriction on that is analogous to the one used when disentangling the efects of AMTR relative to ATR shocks. This implies that they exactly identify both the AMTR shocks to the top 1 percent and the AMTR shocks to the bottom 99 percent by imposing a zero restriction on the systematic component of tax policies. Second, we show that in this case—even when one decides to ignore the robustness issues raised above and focuses on two orderings—the IRFs are more sensitive to the ordering of the variables. As in Mertens and Montiel-Olea’s (2018) second finding, when reporting results for the shocks to the AMTR of the top 1 percent, they choose a diferent ordering than when reporting results for the AMTR shocks to the bottom 99 percent. We find this confusing since in fact their results are not comparable because they come from two diferent identification schemes. In any case, we again substitute the zero restriction for a set of less restrictive sign restrictions to be described below, and we show that our results are very much in line with the results that Mertens and Montiel-Olea (2018) seem to favor when analyzing their results.

In this application we use the dataset built by Mertens and Montiel-Olea (2018), which we downloaded from Karel Mertens’s website.17 A detailed description of the dataset can be found in Appendix A of Mertens and Montiel-Olea (2018).

6.1 Macroeconomic Responses to Marginal Tax Rates

Let’s begin by revisiting the first finding. In their benchmark specification, Mertens and Montiel-Olea (2018) use yearly data from 1946 through 2012 to estimate a proxy-SVAR including nine endogenous variables, two exogenous variables, and one proxy for AMTR shocks. The endogenous variables are the negative of log net-of-tax rate, log reported income, log real GDP per tax unit, the unemployment rate, the log real stock market index, inflation, the federal funds rate, log real government spending per tax unit, and the change in log real federal government debt per tax unit.18 The exogenous variables are dummy variables for the years 1949 and 2008. The proxy (which we call the AMTR proxy) is a collection of instances of variation in marginal tax rates that the authors reasonably consider to be contemporaneously exogenous changes in the AMTR.19 Accordingly, the identification of the AMTR shock is achieved by assuming that the proxy is only correlated with the AMTR.20 Following Mertens and Montiel-Olea (2018), the SVAR features two lags and a constant term. Altogether, in this application T = 65, n = 9, k = 1, p = 2, ˜e = 2, and

17Link to the dataset: https://karelmertenscom.files.wordpress.com/2018/01/data_mmo.xlsx.
18Net-of-tax rate is defined as 1 minus the AMTR.
Figura
Figura
Figura

Figure 2: IRFs to a positive AMTR shock (rate cut). The solid curves represent the point-wise posterior medians, and the shaded areas represent the 68 percent equal-tailed point-wise probability bands. The figure is based on 10,000 independent efective draws obtained using Algorithm 2.

Figure 2: IRFs to a positive AMTR shock (rate cut). The solid curves represent the point-wise posterior medians, and the shaded areas represent the 68 percent equal-tailed point-wise probability bands. The figure is based on 10,000 independent efective draws obtained using Algorithm 2.

We set and to characterize our prior over the proxy-SVAR structural parameters, and we set and to characterize our proposal over the orthogonal-triangular-block parameterization.21 We use Algorithm 2 to obtain an efective sample size of

19The net-of-tax rate is based on Barro and Redlick (2011); for additional details on the construction of the endogenous variables and the proxy, we refer the reader to Mertens and Montiel-Olea’s (2018) paper.
20The sign of the AMTR shock is pinned-down by assuming that the IRF of AMTR is negative in response to a negative AMTR shock.
21In addition, we impose a lower bound on the reliability of the instruments tailored to maximize the eficiency of the importance sampler. In particular, we assume that the minimum eigenvalue of the reliability matrix implied by the posterior distribution of the structural parameters must be greater or equal than 0.2. This implies that at least 20 percent of the variance of any linear combination of the instruments must be related to the underlying structural shocks of interest. See Mertens and Ravn (2013) and Gleser (1992) for details on the reliability matrix. Without this threshold our algorithm becomes ineficient in the subsequent variant of this application, i.e. Section 6.2. For consistency, we use the same lower bound throughout Section 6 although it only becomes crucial for the eficiency of the algorithm in Section 6.2. We think that this issue is due to the fact that

10,000 draws of proxy-SVAR structural parameters satisfying the exogeneity restrictions.

Figure 2 shows the point-wise median and the 68 percent equal-tailed point-wise probability bands for the IRFs of the key variables of interest to a positive one standard deviation AMTR shock. The shock lowers the tax rate for at least 4 years. Clearly, the positive and sizable IRFs of real GDP and the negative and sizable IRFs of the unemployment rate coincide with a positive and sizable response of income. Therefore, our results clearly align with those reported in Figure 5 of Mertens and Montiel-Olea (2018).

In Appendix A.9 we show that Koopman, Shephard and Creal’s (2009) tests validate the finite variance requirement for the reliable use of Algorithm 2.

6.2 Average versus Marginal Tax Rates

The results shown in the previous section imply that AMTR shocks have sizable efects on economic activity. For example, a negative one standard deviation AMTR shock boosts real GDP by about 0.7 percent after one year. Next, Mertens and Montiel-Olea (2018) compare the efects of these shocks with those of ATR shocks in order to understand whether tax policy operates through substitution efects associated with AMTR shocks or through income efects associated with ATR shocks, which ultimately can provide new insights for assessing the economic consequences of tax reforms.

To address this issue, Mertens and Montiel-Olea (2018) expand the SVAR used in Section 6.1 by adding the log ATR as an endogenous variable.22 They identify AMTR and ATR shocks using two proxies. Analogously to the case of the AMTR proxy, the new proxy (which we call the ATR proxy) is a collection of instances of variation in ATRs that the authors reasonably consider to be contemporaneously exogenous changes in the ATR. The identification of the AMTR and ATR shocks is achieved in two steps. First, Mertens and Montiel-Olea (2018) assume that the proxies are only correlated with the tax rate shocks. As shown in Proposition 2, this would only separate the tax rates shocks from the rest of the structural shocks but would not allow the researcher to identify the two tax rates shocks individually. Thus, they also use a particular Cholesky decomposition of a sub-matrix of to isolate one shock from the other. In Appendix A.10 we show how such a Cholesky decomposition imposes a zero restriction on (and thereby on which implies that the AMTR cannot react contemporaneously to the ATR if the AMTR is ordered first, while the ATR cannot react contemporaneously to the AMTR if ATR ordered first. Evidently, this additional zero restriction is a restriction on the systematic component of tax policy and it will exactly identify both the AMTR and

the two tax instruments to be used in this application have many zeros.
22ATR is defined as total revenue and contributions as a ratio of the Piketty and Saez (2003) measure of aggregate market income.

ATR shocks. When analyzing Mertens and Montiel-Olea’s (2018) results contrasting the efects of AMTR and ATR we will order AMTR first because that is the identification scheme under which they conduct such comparison.

Thus, in this application we have T = 65, n = 10, e˜ = 2, and . We set ν = n˜, and to characterize our prior over the proxy-SVAR structural parameters, and we set νˆ = ν˜, 2 and to characterize our proposal over the orthogonal-triangular-block parameterization. We choose to maximize the eficiency of the importance sampler.23 We use Algorithm 4 to obtain an efective sample size of 10,000 draws.

Figure 3a shows the point-wise median and the 68 percent equal-tailed point-wise probability bands for the IRFs of the key variables of interest to a negative one standard deviation AMTR (gray) and ATR (red) shock when the zero restriction is used. Essentially, this panel replicates Mertens and Montiel-Olea’s (2018) exercise from a Bayesian perspective. As the reader can see, the panel closely resembles the IRFs reported in Panels (A) and (C) of Figure 10 of Mertens and Montiel-Olea (2018) and it is very easy to conclude that income, GDP, and the unemployment rate only react to AMTR shocks. This figure justifies the following claims: “There is, on the other hand, no evidence for any efect on incomes when average tax rates decline but marginal rates do not” (Mertens and Montiel-Olea, 2018, page 3), and “The main finding is that, in sharp contrast to the results for marginal tax rate changes after controlling for average tax rates, there is no evidence that income responds strongly to average tax rate changes once marginal rate changes are controlled for. The point estimates are in fact slightly negative, although they are not statistically significant at any horizon.” (Mertens and Montiel-Olea, 2018, page 35).

But one may think that the zero restriction imposed by Mertens and Montiel-Olea (2018) is too restrictive. We now check how robust the results are to this zero restriction by substituting it with a set of less restrictive sign restrictions. As we will see below, the results obtained using these sign restrictions suggest caution while reading Mertens and Montiel-Olea’s (2018) findings. But, before discussing them in more detail, let us describe our sign restrictions.

Sign Restrictions for Identifying AMTR vs ATR Shocks. (i) The proxy for the AMTR shock is positively correlated with the AMTR shocks; (ii) The proxy for the ATR shock is positively correlated with the ATR shocks; (iii) The covariance between the AMTR shock and the AMTR proxy is bigger than the covariance between the ATR shock and the AMTR proxy; and (iv) The covariance between the ATR shock and the ATR proxy is bigger than the covariance between the AMTR shock and the ATR proxy.

23If we set the algorithm becomes very ineficient. The basic description of the approach used for the selection of is described in Appendix A.5.

Figura
Figura
Figura

(a) Mertens and Montiel-Olea (2018)

(a) Mertens and Montiel-Olea (2018)
Figura
Figura
Figura

(b) Sign Restrictions Figure 3: IRFs to a one standard deviation to the AMTR and ATR shocks. The solid curves (blue for the AMTR shock and red for the ATR shock) represent the point-wise posterior medians, and the shaded areas (gray for the AMTR shock and red for the ATR shock) represent the 68 percent equal-tailed point-wise probability bands. The figure is based on 10,000 independent efective draws obtained using Algorithm 4.

(b) Sign Restrictions Figure 3: IRFs to a one standard deviation to the AMTR and ATR shocks. The solid curves (blue for the AMTR shock and red for the ATR shock) represent the point-wise posterior medians, and the shaded areas (gray for the AMTR shock and red for the ATR shock) represent the 68 percent equal-tailed point-wise probability bands. The figure is based on 10,000 independent efective draws obtained using Algorithm 4.

The implementation of our sign restrictions needs a function and matrices very similar to the ones described in Section 5. In the interest of space, we do not describe them. It is important to emphasize that, while Mertens and Montiel-Olea (2018) exactly identify the two shocks, our sign restrictions set identify them

Figure 3b uses the set of signs restrictions described above instead of the zero restriction implied by the Cholesky decomposition. The dynamic responses of income, real GDP, and the unemployment rate change substantially relative to those shown in Figure 3a. In Figure 3b the 68 percent equal-tailed point-wise probability bands for the IRF of income are significantly above zero for both AMTR and ATR shocks. When looking at real GDP and the unemployment rate we also observe diferences in the results. The 68 percent equal-tailed point-wise probability bands for the IRF of real GDP to both shocks are similar when sign restrictions are used. Turning to the unemployment rate, the IRF of the unemployment rate to an ATR shock is now mostly negative. The diferences in the results are confirmed when analyzing Table 1.

Table 1: Short-run elasticities of income (Inc) and real GDP to tax shocks

Mertens and Montiel-Olea (2018)
Ratio of IRFs $Inc_{t+1}/AMTR_t$ $Inc_{t+1}/ATR_t$ $GDP_{t+1}/AMTR_t$ $GDP_{t+1}/ATR_t$
Median1.520.120.850.20
68% Prob. Interval[0.97; 2.15][-0.46; 1.06][0.48; 1.28][-0.20; 0.83]
Sign Restrictions
Ratio of IRFs $Inc_{t+1}/AMTR_t$ $Inc_{t+1}/ATR_t$ $GDP_{t+1}/AMTR_t$ $GDP_{t+1}/ATR_t$
Median1.540.520.790.38
68% Prob. Interval[0.92; 2.32][0.16; 1.15][0.37; 1.30][0.13; 0.83]

Note: The entries in the table denote the posterior moments of the ratio between the IRF of income (Inc) and real GDP one period after the shock and the IRF of the AMTR and ATR on impact following an AMTR and ATR shock, respectively. See the main text for details. The table is based on the same 10,000 independent efective draws obtained using Algorithm 4 used in Figure 3.

This table shows the short-run elasticities of income and real GDP to AMTR and ATR shocks when the zero restrictions are used and when they are substituted by the set of sign restrictions. The short-run elasticities are measured by the ratio between the IRF of income (real GDP) one period after the shock and the impact IRF of the AMTR (ATR) to an AMTR (ATR) shock. Accordingly, these elasticities can be interpreted as the percent change in income and real GDP following an AMTR or ATR shock that decreases the corresponding tax rate by about 1 percentage point. The reader can see that, when using the Cholesky decomposition, the 68 percent posterior probability intervals for the short-run elasticities of income and real

GDP to ATR shocks include negative numbers and that the posterior median is quite low when compared to the AMTR case. That is not the case when we use the set of sign restrictions instead. Although lower than the ones corresponding to the AMTR shocks, the short-run elasticities of income and real GDP to ATR shocks are clearly positive.

Comparing Figures 3a and 3b and reading the results in Table 1, it becomes clear that when using our less restrictive identification scheme it is very dificult to claim that “There is, on the other hand, no evidence for any efect on incomes when ATRs decline but marginal rates do not” (Mertens and Montiel-Olea, 2018, page 2) or “there is no evidence that income responds strongly to ATR changes once marginal rate changes are controlled for.” (Mertens and Montiel-Olea, 2018, page 35). It is true that there may be other restrictions consistent with the results in Mertens and Montiel-Olea (2018). Nevertheless, we think that our results make apparent that further analysis is needed before ruling out the income efects of exogenous changes in average tax cut rates.

In Appendix A.11 we show that Koopman, Shephard and Creal’s (2009) tests validate the finite variance requirement for the reliable use of Algorithm 4.

6.3 Marginal Rates Cuts for the Top and Bottom of the Income Distribution

We now turn to the third result highlighted by Mertens and Montiel-Olea (2018). According to their findings, negative AMTR shocks for taxpayers at the top 1 percent of the income distribution have positive short-run efects on income and economic activity but zero long-run efects (i.e. 4 to 5 years after the shock); in contrast, negative AMTR shocks for tax payers at the bottom 99 percent of the income distribution have zero short-run efects on income and economic activity but positive long-run efects.

As was the case when addressing the efects of AMTR relative to ATR rate cuts, Mertens and Montiel-Olea (2018) modify the SVAR used in Section 6.1 by introducing additional variables germane to the question under study. Specifically, they replace the negative of the aggregate log net-of-tax rate with the negative of the log net-of-tax rate for the top 1 percent and bottom 99 percent of the income distribution, and the aggregate log income level with the log income levels for the top 1 percent and bottom 99 percent of the income distribution. In addition, they modify the reduced-form specification by including a linear and a quadratic trend to capture longer trends in income inequality following Saez (2004) and Saez, Slemrod, and Giertz (2012).

Mertens and Montiel-Olea (2018) parse out the efects of average marginal personal tax rates shocks at the top 1 percent of the income distribution relative to shocks at the bottom 99 percent by using two newly built disaggregated measures of exogenous variation in tax rates across the income distribution as proxies for top and bottom marginal tax rate shocks. As was the case in Section 6.2, Mertens and Montiel-Olea (2018) exactly identify both tax rate shocks by combining the exogeneity restrictions with an additional zero restriction on (and hence on imposed by means of a Cholesky decomposition of a sub-matrix of . This implies that the AMTR of the top 1 percent cannot react contemporaneously to the AMTR of the bottom 99 percent (or vice-versa) depending on the ordering. Consequently, this additional zero restriction is again a restriction on the systematic component of tax policy. When assessing the efects of an AMTR shock for the top 1 percent they choose the order in which the AMTR for the bottom 99 percent cannot react contemporaneously to the AMTR for the top 1 percent. We will refer to this identification scheme as Case I. When assessing the efects of an AMTR for the bottom 99 percent, Mertens and Montiel-Olea (2018) choose the order in which the AMTR for the top 1 percent cannot react contemporaneously to the AMTR for the bottom 99 percent. We will refer to this identification scheme as Case II.

In this application , n = 11, k = 2, p = 2, e˜ = 4, and . We set 2 and to characterize our prior over the proxy-SVAR structural parameters, and we set , and to characterize our proposal over the orthogonal-triangular-block parameterization. We choose to maximize the eficiency of the importance sampler.24 We use Algorithm 4 to obtain an efective sample size of 10,000.

Figure 4 shows the IRFs to a one standard deviation shock to the top 1 percent and bottom 99 percent personal income marginal tax rate, respectively, for Case I. The solid-dotted blue line and the gray area show the point-wise posterior medians and the 68 percent equal-tailed point-wise posterior probability bands to an AMTR shock to the top 1 percent. These IRFs replicate the results in Figure 11 of Mertens and Montiel-Olea (2018). The solid red line and the red area show the point-wise posterior medians and the 68 percent equal-tailed posterior probability bands to an AMTR shock to the bottom 99 percent; these IRFs are not reported by Mertens and Montiel-Olea (2018).

Figure 5 shows the IRFs to a one standard deviation shock to the top 1 percent and bottom 99 percent personal income marginal rate, respectively, for Case II. The red solid line and the red area show the point-wise posterior medians and the 68 percent equal-tailed posterior probability bands to an AMTR shock to the bottom 99 percent. These IRFs correspond to the results in Figure 12 of Mertens and Montiel-Olea (2018). By comparing the latter figure with our results, it becomes evident that we find less support for Mertens and Montiel-Olea’s (2018) conclusions under this identification scheme. The solid-dotted blue line and the gray area show the point-wise posterior medians and the 68 percent equal-tailed point-wise posterior probability bands to an AMTR shock to the top 1 percent; these IRFs are not reported by Mertens and Montiel-Olea (2018).

24If we set the algorithm becomes very ineficient. The basic description of the approach used for the selection of is described in Appendix A.5.
Figura
Figura
Figura
Figura
Figura

Figure 4: IRFs to a one standard deviation shock to the top 1 percent and bottom 99 percent personal income marginal rate, respectively, identified with the Case I scheme. The solid-dotted blue line and the gray area show the point-wise posterior medians and the 68 percent equal-tailed posterior probability bands to an AMTR shock to the top 1 percent. The solid red line and the red area show the point-wise posterior medians and the 68 percent equal-tailed posterior probability bands to an AMTR shock to the bottom 99 percent. The figure is based on 10,000 independent draws obtained using Algorithm 4.

Figure 4: IRFs to a one standard deviation shock to the top 1 percent and bottom 99 percent personal income marginal rate, respectively, identified with the Case I scheme. The solid-dotted blue line and the gray area show the point-wise posterior medians and the 68 percent equal-tailed posterior probability bands to an AMTR shock to the top 1 percent. The solid red line and the red area show the point-wise posterior medians and the 68 percent equal-tailed posterior probability bands to an AMTR shock to the bottom 99 percent. The figure is based on 10,000 independent draws obtained using Algorithm 4.
Figura
Figura
Figura
Figura
Figura

Figure 5: IRFs to a one standard deviation shock to the top 1 percent and bottom 99 percent personal income marginal rate, respectively, identified with the Case II scheme. The solid-dotted blue line and the gray area show the point-wise posterior medians and the 68 percent equal-tailed posterior probability bands to an AMTR shock to the top 1 percent. The solid red line and the red area show the point-wise posterior medians and the 68 percent equal-tailed posterior probability bands to an AMTR shock to the bottom 99 percent. The figure is based on 10,000 independent draws obtained using Algorithm 4.

Figure 5: IRFs to a one standard deviation shock to the top 1 percent and bottom 99 percent personal income marginal rate, respectively, identified with the Case II scheme. The solid-dotted blue line and the gray area show the point-wise posterior medians and the 68 percent equal-tailed posterior probability bands to an AMTR shock to the top 1 percent. The solid red line and the red area show the point-wise posterior medians and the 68 percent equal-tailed posterior probability bands to an AMTR shock to the bottom 99 percent. The figure is based on 10,000 independent draws obtained using Algorithm 4.

As can be seen, the IRFs to a one standard deviation shock to the top 1 percent and bottom 99 percent marginal tax rates shown in Figure 4 endorse the view put forward by Mertens and Montiel-Olea (2018).

Figura
Figura
Figura
Figura
Figura

Figure 6: IRFs to a one standard deviation to the top 1 percent (gray) and bottom 99 percent (red) personal income marginal rate. Instruments plus sign restrictions for identifying AMTR shocks to the top 1 percent and bottom 99 percent of the income distribution. The solid curves represent the point-wise posterior medians, and the shaded areas represent the 68 percent equal-tailed point-wise probability bands. The figure is based on 10,000 independent draws obtained using Algorithm 4.

Figure 6: IRFs to a one standard deviation to the top 1 percent (gray) and bottom 99 percent (red) personal income marginal rate. Instruments plus sign restrictions for identifying AMTR shocks to the top 1 percent and bottom 99 percent of the income distribution. The solid curves represent the point-wise posterior medians, and the shaded areas represent the 68 percent equal-tailed point-wise probability bands. The figure is based on 10,000 independent draws obtained using Algorithm 4.

Nevertheless, Figure 5 makes clear that the results on the efects of tax rate cuts shocks to the bottom 99 percent are not robust to using Case II. In particular, even though under Case I the 68 percent posterior probability IRFs for real GDP and the unemployment rate are above and below zero respectively, the 68 percent probability bands for these IRFs do contain zero under Case II. This implies that, while we are able to replicate Mertens and Montiel-Olea’s (2018) findings under one ordering, we are not able to replicate them under the alternative order. The fact that a potentially influential result hinges on the ordering of an additional zero restriction makes it less appealing. Furthermore, note that under Case II the IRFs of income for the top 1 percent as well as the IRFs for the AMTR for the top 1 percent to a one standard deviation shock to the bottom 99 percent marginal tax rate are not well identified.

Interestingly, next we will show that the lack of robustness and identification vanishes once we replace Mertens and Montiel-Olea’s (2018) zero restriction on tax policy with less restrictive sign restrictions analogous to those used to parse out AMTR from ATR shocks. More specifically, when using the less restrictive sign restrictions described below we are able to confirm Mertens and Montiel-Olea’s (2018) conclusions.

Sign Restrictions for Identifying AMTR Shocks to the Top 1 and Bottom 99 percent. (i) The proxy for the AMTR shock to the top 1 percent is positively correlated with the AMTR shock to the top 1 percent; (ii) The proxy for the AMTR shock to the bottom 99 percent is positively correlated with the AMTR shock to the bottom 99 percent; (iii) The covariance between the AMTR shock to the top 1 percent and the proxy for the AMTR shock to the top 1 percent is bigger than the covariance between the AMTR shock to the bottom 99 percent and the proxy for the AMTR shock to the top 1 percent; and (iv) The covariance between the AMTR shock to the bottom 99 percent and the proxy for the AMTR shock to the bottom 99 percent is bigger than the covariance between the AMTR shock to the top 1 percent and the proxy for the AMTR shock to the bottom 99 percent.

Analogously to the case in Section 6.2, the implementation of our sign restrictions needs a function and matrices very similar to the ones described in Section 5. In the interest of space, we do not describe them.

Figure 6 shows the IRFs to a one standard deviation AMTR shock to the top 1 percent and bottom 99 percent, respectively, identified with our sign restrictions. The red solid line and the red area show the point-wise posterior medians and the 68 percent equal-tailed posterior probability bands to an AMTR shock to the bottom 99 percent. The solid-dotted blue line and the gray area show the point-wise posterior median and the 68 percent equal-tailed posterior probability bands to an AMTR shock to the top 1 percent. It is clear from the figure that Mertens and Montiel-Olea’s (2018) conclusions regarding the efects of tax cut rates shocks at the top and bottom of the distribution can be supported on the grounds of the exogeneity restrictions and our sign restrictions.

In Appendix A.12 we show that Koopman, Shephard and Creal’s (2009) tests validate the finite variance requirement for the reliable use of Algorithm 4.

7 Conclusion

This paper develops eficient algorithms to independently draw from the normal-generalized-normal family of conjugate posterior distributions over the structural parameterization of a Bayesian proxy-SVAR. In addition, our approach expands the type of identification schemes that can be considered under the frequentist paradigm, e.g., Montiel-Olea, Stock and Watson (2016). More specifically, influential papers using the frequentist paradigm rely on additional and often questionable zero restrictions when more than one instrument is used to identify more than one structural shock. In contrast, our Bayesian approach allows researchers to consider less restrictive identification schemes.

A Appendix

A.1 Proof of Proposition 1

Proof. It is well known that in the class of Gaussian linear models considered in this paper and are observationally equivalent if and only if they have the same reduced-form parameterization. They will have the same reduced-form parameterization if and only if

\[\left(\bar {\boldsymbol {A}} _ {0} \bar {\boldsymbol {A}} _ {0} ^ {\prime}\right) ^ {- 1} = \left(\hat {\boldsymbol {A}} _ {0} \hat {\boldsymbol {A}} _ {0} ^ {\prime}\right) ^ {- 1}\tag{8}\]

\[\bar {\boldsymbol {A}} _ {+} \bar {\boldsymbol {A}} _ {0} ^ {- 1} = \hat {\boldsymbol {A}} _ {+} \hat {\boldsymbol {A}} _ {0} ^ {- 1}.\tag{9}\]

First, assume that and are observationally equivalent so that Equations (8) and (9) hold. We show that there is such that and . Let . It follows directly from the definition of that , and Equation (9) implies . We now show that . It follows from Equation (8) that . Hence, is orthogonal. Because the lower left-hand block of both and are zero, the same will be true for both and . Let

\[\boldsymbol {Q} = \left[ \begin{array}{c c} \boldsymbol {Q} _ {1} & \boldsymbol {W} \\ \boldsymbol {0} _ {k \times n} & \boldsymbol {Q} _ {2} \end{array} \right],\]

where is is , and W is . Because , it must be the case that and . But , if and only if . So

Now, assume that and for some matrix . It is easy to see by direct substitution that Equations (8) and (9) hold, which implies that and are observationally equivalent. □

A.2 The Function

In Steps 1, 2, and 3 of Algorithm 1 draws of were obtained.25 For each draw, Steps 4 and 5 of Algorithm 1 recursively defined matrices and , though the were not uniquely defined. Using the , we can define a mapping g of to , where . In this appendix we show that there is an open set such that g can be uniquely defined over U so that it is continuously diferentiable. Furthermore, U will contain almost all draws of .

25In this appendix it will always be the case that and when i = 1 and when

To define the we need reference matrices that are of full column rank, where

\[m _ {i} = \left\{ \begin{array}{l l} n & \text {if} i = 1 \\ k & \text {if} i = 2 \end{array} \right. \quad \text {and} \quad n _ {i, j} = \left\{ \begin{array}{c c} n + 1 - j - k & \text {if} i = 1 \text {and} j \leq n - k \\ n + 1 - j & \text {if} i = 1 \text {and} j > n - k \\ k + 1 - j & \text {if} i = 2 \end{array} \right..\]

We recursively define , and . Let in , define

\[\tilde {\boldsymbol {M}} _ {i, j} (\tilde {\boldsymbol {\Lambda}} _ {0}, \tilde {\boldsymbol {\Lambda}} _ {+}, \boldsymbol {w} _ {i, j}) = \left\{ \begin{array}{c c} \left[ \begin{array}{c c c c} \boldsymbol {K} _ {i, 1} \boldsymbol {w} _ {i, 1} & \dots & \boldsymbol {K} _ {i, j - 1} \boldsymbol {w} _ {i, j - 1} & \boldsymbol {G} (\tilde {\boldsymbol {\Lambda}} _ {0}) ^ {\prime} \\ & & & \boldsymbol {W} _ {i, j} \end{array} \right] & i = 1 \text {and} j \leq n - k \\ \left[ \begin{array}{c c c c} \boldsymbol {K} _ {i, 1} \boldsymbol {w} _ {i, 1} & \dots & \boldsymbol {K} _ {i, j - 1} \boldsymbol {w} _ {i, j - 1} & \boldsymbol {W} _ {i, j} \end{array} \right] & \text {otherwise} \end{array} \right..\]

Next, let

\[U _ {i, j} = \left\{(\tilde {\boldsymbol {\Lambda}} _ {0}, \tilde {\boldsymbol {\Lambda}} _ {+}, \boldsymbol {w} _ {i, j}) \in U _ {i, j - 1} | \det (\tilde {\boldsymbol {M}} _ {i, j} (\tilde {\boldsymbol {\Lambda}} _ {0}, \tilde {\boldsymbol {\Lambda}} _ {+}, \boldsymbol {w} _ {i, j})) \neq 0 \right\}.\]

Because the determinant function is continuous, will be open. Let

\[\tilde {\boldsymbol {M}} _ {i, j} (\tilde {\boldsymbol {\Lambda}} _ {0}, \tilde {\boldsymbol {\Lambda}} _ {+}, \boldsymbol {w} _ {i, j}) = \boldsymbol {Q} _ {i, j} (\tilde {\boldsymbol {\Lambda}} _ {0}, \tilde {\boldsymbol {\Lambda}} _ {+}, \boldsymbol {w} _ {i, j}) \boldsymbol {R} _ {i, j} (\tilde {\boldsymbol {\Lambda}} _ {0}, \tilde {\boldsymbol {\Lambda}} _ {+}, \boldsymbol {w} _ {i, j})\]

be the QR-decomposition of normalized so that the diagonal of is positive and define to be the last columns of . The QR-decomposition of a square non-singular matrix, normalized so that the diagonal of the triangular component is positive, is a continuously diferentiable function. Thus, is defined and continuously diferentiable over . Let U be the intersection of all the . Note that U is open and can be defined over U so that it is continuously diferentiable.

All that remains to be shown is that U contains almost all draws of . It sufices to show that , the complement of in , is of measure zero. We will show that is a union of manifolds of strictly lower dimension and thus is of measure zero. For each ` with and each set of ` columns of , consider the set of all such that the given set of ` columns are linearly independent and the remaining columns are in the span of the ` linearly independent columns. An element of will be in at least one such set. Furthermore, sets of these forms will be manifolds of strictly lower dimension.26 This completes the proof.

26For any set of ` columns of , we can project each of the remaining columns onto perpendicular componen of the span of the ` selected columns. As long as the selected columns are linearly independent, this mapping will be continuously

A.3 Proof of Proposition 2

Proof. By Proposition 1, the proxy-SVAR structural parameters and are observationally equivalent if and only if there exists such that and . So to prove Proposition 2, it sufices to prove that if and satisfy the exogeneity restrictions and the relevance condition, then will satisfy the exogeneity restrictions and relevance condition if and only if . Recall from Sections 3.2 and 3.3, the proxy-SVAR parameters will satisfy the relevance condition if and only if is of full row rank and will satisfy the exogeneity restrictions if and only if 2 where

We will maintain the assumption that and satisfy the exogeneity restrictions and the relevance condition. First assume that , so that , where (k), and . Since and satisfy the relevance condition and the exogeneity restriction, so will

Now assuming that the parameters satisfy the relevance condition and the exogeneity restriction, we show that . Because it is of the form , where and . Let

\[\boldsymbol {Q} _ {1} = \left[ \begin{array}{c c} \boldsymbol {Q} _ {1 1} & \boldsymbol {Q} _ {1 2} \\ \boldsymbol {Q} _ {2 1} & \boldsymbol {Q} _ {2 2} \end{array} \right],\]

where is is is , and is . If is zero, then it will follow that . Because satisfies both the relevance condition and the exogeneity restrictions, the columns of form a basis for the null space of . Because both and satisfy the exogeneity restrictions, then

\[\mathbf {0} _ {k \times n - k} = J ((\hat {\boldsymbol {A}} _ {0} \boldsymbol {Q}) ^ {- 1}) ^ {\prime} \boldsymbol {L} ^ {\prime} \boldsymbol {K} ^ {\prime} = J (\hat {\boldsymbol {A}} _ {0} ^ {- 1}) ^ {\prime} \boldsymbol {L} ^ {\prime} \left(\boldsymbol {K} ^ {\prime} \boldsymbol {Q} _ {1 1} + \left[ \begin{array}{c} \mathbf {0} _ {n - k, n - k} \\ \boldsymbol {Q} _ {2 1} \end{array} \right]\right) = \boldsymbol {J} (\hat {\boldsymbol {A}} _ {0} ^ {- 1}) ^ {\prime} \boldsymbol {L} ^ {\prime} \left[ \begin{array}{c} \mathbf {0} _ {n - k, n - k} \\ \boldsymbol {Q} _ {2 1} \end{array} \right].\]

Thus must be zero as desired.

A.4 Gibbs Sampler

In Waggoner and Zha (2003), a Gibbs sampler is described for sampling from a posterior distribution of a structural VAR over a certain class of normal priors and subject to a certain class of linear non-cross equation restrictions. In that paper, the restrictions are described in terms of free parameters: In particular, if and denote the columns of and , respectively, then it is assumed that the and that satisfy the restrictions are of the form

diferentiable and the remaining columns will be in the span of the selected columns if and only if this mapping is zero. This mapping along with standard results from multivariate calculus is enough to imply that the sets defined above are manifolds of strictly lower dimension.

\[\tilde {\boldsymbol {\lambda}} _ {0, j} = \boldsymbol {U} _ {j} \boldsymbol {\gamma} _ {0, j} \text {and} \tilde {\boldsymbol {\lambda}} _ {+, j} = \boldsymbol {V} _ {j} \boldsymbol {\gamma} _ {+, j},\]

where both and have orthonormal columns for . Because must be upper triangular, can be taken to be the first columns of . Because satisfy the block restrictions, can be taken to be for . When , will be block diagonal with the first blocks equal to the first n columns of and the last block the scalar one.

The Gibbs sampler is described in terms of a non-negative scalar and matrices and , for . In that paper, the goal was to sample from a posterior and so T , and were given in terms of restrictions, prior, and data. Our goal is to sample from a normal-generalized-normal distribution conditional on the above restrictions, so we will describe T , and in terms , and and the above restrictions. The , and must be block diagonal, so assume that , and ).

\[\begin{array}{c} T = \nu - \tilde {n} \\ \boldsymbol {H} _ {j} = (\boldsymbol {V} _ {j} ^ {\prime} \hat {\boldsymbol {\Omega}} _ {j} ^ {- 1} \boldsymbol {V} _ {j}) ^ {- 1} \\ \boldsymbol {P} _ {j} = \boldsymbol {H} _ {j} \boldsymbol {V} _ {j} ^ {\prime} \hat {\boldsymbol {\Omega}} _ {j} ^ {- 1} \hat {\boldsymbol {\Psi}} \boldsymbol {U} _ {j} \\ \boldsymbol {S} _ {j} = \left(\frac {1}{T} (\boldsymbol {U} _ {j} ^ {\prime} \hat {\boldsymbol {\Phi}} \boldsymbol {U} _ {j} + \boldsymbol {U} _ {j} ^ {\prime} \hat {\boldsymbol {\Psi}} ^ {\prime} \hat {\boldsymbol {\Omega}} ^ {- 1} \hat {\boldsymbol {\Psi}} \boldsymbol {U} _ {j} - \boldsymbol {P} _ {j} ^ {\prime} \boldsymbol {H} _ {j} ^ {- 1} \boldsymbol {P} _ {j})\right) ^ {- 1} \end{array}\]

A.5 Proposal Normal-Generalized-Normal Parameters

As mentioned in Section 3, while oftentimes it sufices to choose to be equal to , there are instances in which this can lead to small efective sample sizes in our importance sampler. In such cases we find it useful to tailor the choice of by choosing the value of that minimizes the squared of the diference between the target and the proposal density evaluated at a given number of draws of the posterior distribution over the structural parameterization obtained when is set equal to

27In the applications of our paper that rely on this procedure (i.e. Sections 6.2 and 6.3), performing this optimization based on 1,000 draws of the structural parameters sufices to find an eficient choice of Φ .

A.6 Data Appendix for Section 5

Here we describe the data used in Section 5 in more details. The time series used to construct the endogenous variables used in the proxy-SVAR are:

1. Real Gross Domestic Product, BEA, NIPA table 1.1.6, line 1, billions of chained (2009) dollars, seasonally adjusted at annual rates. Downloaded from https://www.bea.gov.

2. Total Private Employment, BLS, Current Employment Statistics survey (National), series Id CES0500000001, thousands, seasonally adjusted. Downloaded from https://www.bls.gov.

3. Price Index for Gross Domestic Product, BEA, NIPA table 1.1.4, line 1, index 2009=100, seasonally adjusted. Downloaded from https://www.bea.gov.

4. Personal Consumption Expenditures on Non-durable Goods, BEA NIPA table 1.1.5, line 5, billions of dollars, seasonally adjusted at annual rate. Downloaded from https://www.bea.gov.

5. Personal Consumption Expenditures on Services, BEA NIPA table 1.1.5, line 6, billions of dollars, seasonally adjusted at annual rate. Downloaded from https://www.bea.gov.

6. Personal Consumption Expenditures on Durable Goods, BEA NIPA table 1.1.5, line 4, billions of dollars, seasonally adjusted at annual rate. Downloaded from https://www.bea.gov.

7. Fixed Investment in Equipment, BEA NIPA table 1.1.5, line 8, billions of dollars, seasonally adjusted at annual rate. Downloaded from https://www.bea.gov.

8. Real Consumption = (4)+(5) / (3)

9. Real Investment in Equipment = (6)+(7) / (3)

The endogenous variables in the SVAR are series (1), (2), (3), (8), and (9) transformed to percent log diferences.

A.7 Finite Variance Tests of Importance Sampling Weights

We numerically test for the variance of the importance sampler weights to be finite in each of the applications of the paper. In particular we use the Wald, score, and likelihood ratio (LR) tests as described in Koopman, Shephard, and Creal (2009). These tests assume that the importance sampler weights are independent draws from a Pareto distribution characterized by the shape parameter ξ. The null of each one of these tests is

\[H _ {0}: \xi = \frac {1}{2} \text { and } H _ {1}: \xi > \frac {1}{2}\]

because for the Pareto distribution variance does not exist.

We conduct the tests for several thresholds of the importance sampler weights ranging from the largest 50 percent to the largest 1 percent of the importance sampler weights. These thresholds determine the number of importance sampler weights used to implement the tests. The 95 percent critical values for the Wald, score, and LR tests are 1.64, 1.64, and 2.69 respectively.

A.8 Tests for the Analysis in Section 5

Table A.1 shows the value of the tests described above for several thresholds—as shown in the first row of the table—applied to the analysis performed in Section 5. The second row of the table shows the Wald test statistics, the third row shows the score test statistics, and the fourth row shows the LR test statistics. None of the values displayed in Table A.1 exceed the critical values reported above. Hence, these tests indicate tha the importance sampler weights have finite variance.

Table A.1: Wald, Score and Likelihood Ratio Tests for the Analysis in Section 5.

ThresholdLargest 50%Largest 40%Largest 30%Largest 10%Largest 1%
Wald-2.36-2.06-1.70-0.86-0.22
Score-9.34-8.07-6.31-3.77-1.25
LR00000

A.9 Tests for the Analysis in Section 6.1

Table A.2 shows that Koopman, Shephard and Creal’s (2009) tests indicate that the importance weights used in Section 6.1 have finite variance.

Table A.2: Wald, Score and Likelihood Ratio Tests for the Analysis in Section 6.1.

ThresholdLargest 50%Largest 40%Largest 30%Largest 10%Largest 1%
Wald-4979.60-4237.73-3478.55-1716.32-355.17
Score-6.07-5.58-5.10-3.99-1.63
LR00000

A.10 Understanding the Additional Zero Restrictions

The reduced-form parameterization implied by Equation (3) is

\[\pmb {y} _ {t} ^ {\prime} = \pmb {x} _ {t} ^ {\prime} \pmb {B} + \pmb {u} _ {t} ^ {\prime},\]

where , and . Thus, we have that

\[\pmb {u} _ {t} = (\pmb {A} _ {0} ^ {- 1}) ^ {\prime} \pmb {\varepsilon} _ {t},\tag{10}\]

where are the reduced-form innovations. We can rewrite Equation (10) as

\[{\left[ \begin{array}{l} \pmb {u} _ {1 t} \\ \pmb {u} _ {2 t} \end{array} \right]} = {\left[ \begin{array}{l l} \pmb {B} _ {1 1} & \pmb {B} _ {1 2} \\ \pmb {B} _ {2 1} & \pmb {B} _ {2 2} \end{array} \right]} {\left[ \begin{array}{l} \pmb {\varepsilon} _ {1 t} \\ \pmb {\varepsilon} _ {2 t} \end{array} \right]} \mathrm{where} (\pmb {A} _ {0} ^ {- 1}) ^ {\prime} = {\left[ \begin{array}{l l} \pmb {B} _ {1 1} & \pmb {B} _ {1 2} \\ \pmb {B} _ {2 1} & \pmb {B} _ {2 2} \end{array} \right]}\]

and are the k structural shocks correlated with the proxies and are the structural shocks that are not correlated with the proxies.28 is matrix, is a matrix, is a matrix, and is a matrix. Hence, we have

\[\pmb {u} _ {1 t} = \pmb {B} _ {1 1} \pmb {\varepsilon} _ {1 t} + \pmb {B} _ {1 2} \pmb {\varepsilon} _ {2 t}\tag{11}\]

and

\[\pmb {u} _ {2 t} = \pmb {B} _ {2 1} \pmb {\varepsilon} _ {1 t} + \pmb {B} _ {2 2} \pmb {\varepsilon} _ {2 t}.\tag{12}\]

From Equation (12), we have that

\[\pmb {\varepsilon} _ {2 t} = - \pmb {B} _ {2 2} ^ {- 1} \pmb {B} _ {2 1} \pmb {\varepsilon} _ {1 t} + \pmb {B} _ {2 2} ^ {- 1} \pmb {u} _ {2 t}.\tag{13}\]

Then, Equations (11) and (13) imply

\[\pmb {u} _ {1 t} = \pmb {B} _ {1 1} \pmb {\varepsilon} _ {1 t} + \pmb {B} _ {1 2} \left(- \pmb {B} _ {2 2} ^ {- 1} \pmb {B} _ {2 1} \pmb {\varepsilon} _ {1 t} + \pmb {B} _ {2 2} ^ {- 1} \pmb {u} _ {2 t}\right)\tag{14}\]

or

\[\pmb {u} _ {1 t} = \eta \pmb {u} _ {2 t} + \pmb {S} _ {1} \pmb {\varepsilon} _ {1 t},\tag{15}\]

where and . Similarly,

\[\pmb {u} _ {2 t} = \zeta \pmb {u} _ {1 t} + \pmb {S} _ {2} \pmb {\varepsilon} _ {2 t},\tag{16}\]

28In Section 2.2 we correlated the last k structural shocks with the proxies. That was without loss of generality. We change the order here to better match the explanations in Mertens and Ravn (2013).

where and . Equations (15) and (16) replicate Equations (15) and (16) in Mertens and Ravn (2013). Using Equation (15) and (16), we get

\[\boldsymbol {u} _ {t} = \left[ \begin{array}{c c} \boldsymbol {I} _ {k} & - \eta \\ - \zeta & \boldsymbol {I} _ {n - k} \end{array} \right] ^ {- 1} \left[ \begin{array}{c c} \boldsymbol {S} _ {1} & \boldsymbol {0} \\ \boldsymbol {0} & \boldsymbol {S} _ {2} \end{array} \right] \varepsilon_ {t} = \left[ \begin{array}{c c} \boldsymbol {I} _ {k} + \eta \left(\boldsymbol {I} _ {n - k} - \zeta \eta\right) ^ {- 1} \zeta & (\boldsymbol {I} _ {k} - \eta \zeta) ^ {- 1} \eta \\ (\boldsymbol {I} _ {n - k} - \zeta \eta) ^ {- 1} \zeta & \boldsymbol {I} _ {n - k} + \zeta \left(\boldsymbol {I} _ {k} - \eta \zeta\right) ^ {- 1} \eta \end{array} \right] \left[ \begin{array}{c c} \boldsymbol {S} _ {1} & \boldsymbol {0} \\ \boldsymbol {0} & \boldsymbol {S} _ {2} \end{array} \right] \varepsilon_ {t}.\tag{17}\]

Hence, we have that

\[(\boldsymbol {A} _ {0} ^ {- 1}) ^ {\prime} = \left[ \begin{array}{c c} \boldsymbol {I} _ {k} + \eta (\boldsymbol {I} _ {n - k} - \zeta \eta) ^ {- 1} \zeta & (\boldsymbol {I} _ {k} - \eta \zeta) ^ {- 1} \eta \\ (\boldsymbol {I} _ {n - k} - \zeta \eta) ^ {- 1} \zeta & \boldsymbol {I} _ {n - k} + \zeta (\boldsymbol {I} _ {k} - \eta \zeta) ^ {- 1} \eta \end{array} \right] \left[ \begin{array}{c c} \boldsymbol {S} _ {1} & \boldsymbol {0} \\ \boldsymbol {0} & \boldsymbol {S} _ {2} \end{array} \right].\tag{18}\]

By letting the first column of be denoted by we see from Equation (18) that

\[\beta_ {1} = \left[ \begin{array}{c} \boldsymbol {I} _ {k} + \eta (\boldsymbol {I} _ {n - k} - \zeta \eta) ^ {- 1} \zeta \\ (\boldsymbol {I} _ {n - k} - \zeta \eta) ^ {- 1} \zeta \end{array} \right] \boldsymbol {S} _ {1}.\tag{19}\]

Mertens and Ravn (2013) show that for any value of the reduced-form parameters one can solve for and , but not for . In particular, on page 1224 the authors write, “Ideally one would like to identify but this requires arbitrary assumptions on how personal income taxes respond contemporaneously to unanticipated changes in corporate taxes (beyond the indirect contemporaneous endogenous efects through , and vice versa. Fortunately, knowledge of still permits economically meaningful structural responses to any linear combination of tax shocks. We report responses that result from a Cholesky decomposition of , imposing that is lower triangular.” Hence, Mertens and Ravn (2013) seem to claim that imposing that is lower triangular does not imply any additional identification restrictions.

Let’s now see what the implications are of imposing that is lower triangular for the matrix describing the contemporaneous relations among the variables, i.e., . First, letting , note that

\[\boldsymbol {A} _ {0} = \left[ \begin{array}{c c} \boldsymbol {L} _ {1} ^ {\prime} & - \zeta^ {\prime} \boldsymbol {L} _ {1} ^ {\prime} \\ - \eta^ {\prime} \boldsymbol {S} _ {2} ^ {\prime} & \left(\boldsymbol {S} _ {2} ^ {- 1}\right) ^ {\prime} \end{array} \right],\tag{20}\]

where . Since is lower triangular, is also lower triangular. Hence, Equation (20) implies that imposing that is lower triangular is indeed imposing additional zero restrictions in . Thus, imposing that is lower triangular does indeed imply some additional zero identification restrictions.

When identifying marginal and average tax shocks, Mertens and Montiel-Olea (2018) use two proxies to identify two structural shocks. Thus, is an upper triangular matrix of dimension Therefore, we have

\[\boldsymbol {A} _ {0} = \left[ \begin{array}{c c} \left[ \begin{array}{c c} u _ {1 1} & u _ {1 2} \\ 0 & u _ {2 2} \end{array} \right] - \zeta^ {\prime} \left(\boldsymbol {S} _ {2} ^ {- 1}\right) ^ {\prime} \\ - \boldsymbol {\eta} ^ {\prime} \boldsymbol {L} _ {1} ^ {\prime} & \left(\boldsymbol {S} _ {2} ^ {- 1}\right) ^ {\prime} \end{array} \right],\]

which implies that the systematic part of AMTR cannot react contemporaneously to ATR if AMTR is ordered first. If ATR is ordered first, ATR cannot react contemporaneously to AMTR.

A.11 Tests for the Analysis in Section 6.2

Table A.3 presents Koopman, Shephard and Creal’s (2009) tests for the case in which the average and marginal tax rates are identified using the restrictions in Mertens and Montiel-Olea (2018) described in Section 6.2. Table A.3: Wald, Score and Likelihood Ratio Tests for the Analysis in Section 6.2 for the proxy-SVAR identified with proxy and zero restrictions.

ThresholdLargest 50%Largest 40%Largest 30%Largest 10%Largest 1%
Wald-17.50-14.25-11.04-4.10-0.48
Score-10.63-9.13-7.71-3.95-0.92
LR00000

Table A.4 shows that Koopman, Shephard and Creal’s (2009) tests indicate that the importance weights used in Section 6.2 for the proxy-SVAR identified with proxy and sign restrictions have finite variance. Table A.4: Wald, Score and Likelihood Ratio Tests for the Analysis in Section 6.2 for the proxy-SVAR identified with proxy and sign restrictions.

ThresholdLargest 50%Largest 40%Largest 30%Largest 10%Largest 1%
Wald-10.23-9.05-7.43-3.62-0.90
Score-6.18-6.03-5.47-3.81-1.72
LR00000

A.12 Tests for the Analysis in Section 6.3

Table A.5 presents Koopman, Shephard and Creal’s (2009) tests for the case in which the bottom and top tax rates cut shocks are identified using the Case I scheme described in Section 6.3.

Table A.5: Wald, Score and Likelihood Ratio Tests for the Analysis in Section 6.3 for the proxy-SVAR identified using the Case I scheme.

ThresholdLargest 50%Largest 40%Largest 30%Largest 10%Largest 1%
Wald-34.51-29.20-23.77-10.35-1.64
Score-11.94-9.97-8.23-4.15-0.92
LR00000

Table A.6 presents Koopman, Shephard and Creal’s (2009) tests for the case in which the bottom and top tax rates cut shocks are identified using the Case II scheme described in Section 6.3. Table A.6: Wald, Score and Likelihood Ratio Tests for the Analysis in Section 6.3 for the proxy-SVAR identified using the the Case II scheme.

ThresholdLargest 50%Largest 40%Largest 30%Largest 10%Largest 1%
Wald-32.99-27.91-22.22-8.41-0.72
Score-11.58-9.67-7.89-3.55-0.54
LR00000

Table A.7 presents Koopman, Shephard and Creal’s (2009) tests for the case in which the bottom and top tax rates cut shocks are identified using the less restrictive scheme described in Section 6.3. Table A.7: Wald, Score and Likelihood Ratio Tests for the Analysis in Section 6.3 for the proxy-SVAR identified using the less restrictive identification scheme.

ThresholdLargest 50%Largest 40%Largest 30%Largest 10%Largest 1%
Wald-28.67-24.77-20.30-9.17-2.03
Score-10.85-9.28-7.52-3.60-1.09
LR00000

References

  1. Arias, J. E., J. F. Rubio-Ramírez, and D. F. Waggoner (2018). Inference Based on Structural Vector Autoregressions Identified with Sign and Zero restrictions: Theory and Applications. Econometrica 86 (2), 685–720.
  2. Bahaj, S. A. (2014). Systemic Sovereign Risk: Macroeconomic Implications in the Euro Area. Centre For Macroeconomics Working Paper .
  3. Barro, R. J. and C. J. Redlick (2011). Macroeconomic Efects from Government Purchases and Taxes. The Quarterly Journal of Economics 126 (1), 51–102.
  4. Caldara, D. and E. Herbst (2016). Monetary Policy, Real Activity, and Credit Spreads: Evidence from Bayesian Proxy SVARs. IFDP (2016-049), Federal Reserve Board.
  5. Drautzburg, T. (2016). A Narrative Approach to a Fiscal DSGE model. Working Paper, FRB Philadelphia.
  6. Fernald, J. (2014). A Quarterly, Utilization-Adjusted Series on Total Factor Productivity. Working Paper 2012-19, Federal Reserve Bank of San Francisco.
  7. Gertler, M. and P. Karadi (2015). Monetary Policy Surprises, Credit Costs, and Economic Activity. American Economic Journal: Macroeconomics 7 (1), 44–76.
  8. Gleser, L. J. (1992). The Importance of Assessing Measurement Reliability in Multivariate Regression. Journal of the American Statistical Association 87 (419), 696–707.
  9. Koopman, S. J., N. Shephard, and D. Creal (2009). Testing the Assumptions Behind Importance Sampling. Journal of Econometrics 149 (1), 2–11.
  10. Leeper, E. M., C. A. Sims, and T. Zha (1996). What Does Monetary Policy Do? Brookings papers on economic activity 1996 (2), 1–78.
  11. Liu, Z., J. Fernald, and S. Basu (2012). Technology Shocks in a Two-Sector DSGE model. Meeting Paper 1017, Society for Economic Dynamics.
  12. Lunsford, K. G. (2016). Identifying Structural VARs with a Proxy Variable and a Test for a Weak Proxy. Federal Reserve Bank of Cleveland Working Paper 15-28 .
  13. Mertens, K. and J. L. Montiel-Olea (2018). Marginal Tax Rates and Income: New Time Series Evidence. Quarterly Journal of Economics 133 (4), 1803–1884.
  14. Mertens, K. and M. O. Ravn (2013). The Dynamic Efects of Personal and Corporate Income Tax Changes in the United States. American Economic Review 103 (4), 1212–47.
  15. Montiel-Olea, J. L., J. H. Stock, and M. W. Watson (2016). Inference in Structural VARs with External Instruments. Working Paper .
  16. Piketty, T. and E. Saez (2003). Income Inequality in the United states, 1913-1998. Quarterly Journal of Economics 118 (1), 1–39.
  17. Rothenberg, T. J. (1971). Identification in Parametric Models. Econometrica 39, 577–591.

Rubio-Ramírez, J., D. Waggoner, and T. Zha (2010). Structural Vector Autoregressions: Theory of Identification and Algorithms for Inference. Review of Economic Studies 77 (2), 665–696.

Saez, E. (2004). Reported Incomes and Marginal Tax Rates, 1960-2000: Evidence and Policy Implications. Tax Policy and the Economy 18, 117–173.

Saez, E., J. Slemrod, and S. H. Giertz (2012). The Elasticity of Taxable Income with Respect to Marginal Tax Rates: A Critical Review. Journal of Economic Literature 50 (1), 3–50.

Sims, C. A. and T. Zha (1998). Bayesian Methods for Dynamic Multivariate Models. International Economic Review 39 (4), 949–968.

Stock, J. H. (2008). What’s New in Econometrics: Time Series, Lecture 7. Short course lectures, NBER Summer Institute at http: // www. nber. org/ minicourse_ 2008. html .

  1. Stock, J. H. and M. W. Watson (2012). Disentangling the Channels of the 2007-09 Recession. Brookings Papers on Economic Activity: Spring 2012 , 81.
  2. Stock, J. H. and M. W. Watson (2018). Identification and Estimation of Dynamic Causal Efects in Macroeconomics Using External Instruments. The Economic Journal 128 (610), 917–948.

Waggoner, D. F. and T. Zha (2003). A Gibbs Sampler for Structural Vector Autoregressions. Journal of Economic Dynamics and Control 28 (2), 349–366.