‹ Volver a la ficha Doc. dt-1998-06

BENCHMARK PRIORS FOR BAYESIAN MODEL AVERAGING

Carmen Fernández Dept. of Mathematics, University of Bristol, Bristol BS8 1TW, U.K.

Eduardo Ley FEDEA, ES-28001 Madrid, Spain

Mark F. J. Steel Dept. of Economics, University of Edinburgh, Edinburgh EH8 9JY, U.K.

March, 1998 Version: April 2, 1998

Abstract. In contrast to a posterior analysis given a particular sampling model, posterior model probabilities in the context of model uncertainty are typically rather sensitive to the specification of the prior. In particular, “difuse” priors on model-specific parameters can lead to quite unexpected consequences. Here we focus on the practically relevant situation where we need to entertain a (large) number of sampling models and we have (or wish to use) little or no subjective prior information. We aim at providing an “automatic” or “benchmark” prior structure that can be used in such cases.

We focus on the Normal linear regression model with uncertainty in the choice of regressors. We propose a partly noninformative prior structure related to a Natural Conjugate g-prior specification, where the amount of subjective information requested from the user is limited to the choice of a single scalar hyperparameter . The consequences of diferent choices for are examined. We investigate theoretical properties, such as consistency of the implied Bayesian procedure. Links with classical information criteria are provided. In addition, we examine the finite sample implications of several choices of in a simulation study. The use of the algorithm of Madigan and York (1995), combined with eficient coding in Fortran, makes it feasible to conduct large simulations. In addition to posterior criteria, we shall also compare the predictive performance of diferent priors. A classic example concerning the economics of crime will also be provided and contrasted with results in the literature. The main findings of the paper will lead us to propose a “benchmark” prior specification in a linear regression context with model uncertainty.

Keywords. Bayes factors, Markov chain Monte Carlo, Posterior odds, Prior elicitation

JEL Classification System. C11, C15

Address. Mark F. J. Steel, Department of Economics, University of Edinburgh, 50 George Square, Edinburgh EH8 9JY, UK; Tel.: +44-131-650 8352, Fax: +44-131-650 4514, Email: Mark.Steel@ed.ac.uk

The issue of variable selection has permeated the econometrics and statistics literature for decades. An enormous volumeofreferences can becited (only afraction ofwhich is mentioned in this paper), and special issues of the Journal of Econometrics (1981, Vol.16, No.1) and Statistica Sinica (1997, Vol.7, No.2) are merely two examples of the amount of interest the topic of model selection has generated in the literature.

Indeed, the issue is an important one, as we are often faced with a situation where a large number of possible regressors can be used and we do not wish to dilute the (often scant) data information we have by including too many regressors. Bayesian methods provide us with a perfectly coherent and interpretable solution, both for selecting and for combining models, through posterior odds. Unfortunately, the influence of the prior distribution, which is often straightforward to assess for inference given the model, is much harder to identify for posterior model probabilities.

Broadly speaking, we can distinguish three strands of related literature in this context. Firstly, we mention the fundamentally oriented statistics and econometrics literature on prior elicitation and model selection, such as exemplified in Box (1980), Zellner and Siow (1980), Draper (1995) and Phillips (1995) and the discussions of these papers. Secondly, there is the recent statistics literature on computational aspects. Markov chain Monte Carlo methods are proposed in George and McCulloch (1993), Green (1995), Madigan and York (1995), Geweke (1996) and Raftery, Madigan and Hoeting (1997), while Laplace approximations arefound in Gelfand and Dey (1995) and Raftery (1996). Finally, thereexists alargeliteratureon information criteria, mostly in the context of time series, see Hannan and Quinn (1979), Akaike (1981), Atkinson (1981), Chow (1981). This paper provides a unifying framework in which these three areas ofresearch will be discussed.

In line with the bulk of the literature, the context of this paper will be that of Normal linear regression with uncertainty in the choice of regressors. We present a prior structure that can reasonably be used in cases wherewehave(orwishto use) littlepriorinformation, partly based on improperpriors forparameters that are common to all models, and partly on a g-prior structure as in Zellner (1986). The prior is not in the natural-conjugate class, but is such that marginal likelihoods can still be computed analytically. This allows for a simple treatment of potentially very large model spaces through Markov chain Monte Carlo model composition as introduced in Madigan and York (1995). In contrast to some ofthe priors proposed in the literature, the prior we propose does not violate the rules ofprobability as it avoids dependence on the values ofthe response variable. The only hyperparameter left to elicit in our prior is a scalar . Theoretical properties, such as consistency ofposterior model probabilities, are linked to functional dependencies of on sample size and the number ofregressors in the model considered. In addition, we conduct an empirical investigation through simulation. This will allow us to suggest specific choices for to the applied user.

As we have conducted a large simulation study, efficient coding was required. This code (in Fortran-77) has been made publicly available on the World Wide Web, and will allow researchers in various (applied) fields to use techniques on large empirically relevant problems at very modest computing costs. In addition, we present the researcher with a simple diagnostic to assess whether the sampler that generates a Monte Carlo Markov chain over model space has converged.

Section 1 introduces the Bayesian model and the practice of Bayesian model averaging. The prior structure is explained in detail in Section 2, where expressions for Bayes factors are also given. Asymptotic consistency of the latter is studied in Section 3. The setup of the empirical simulation experiment is described in Section 4, while results are provided in the next section. An illustrative example using the economic model ofcrime from Ehrlich (1973, 1975) concludes the paper.

1. THE MODEL AND BAYESIAN MODEL AVERAGING

We consider n independent replications from a linear regression model with an intercept, say α, and k possible regression coefficients grouped in a k-dimensional vector . We denote by Z the corresponding design matrix and we assume that , where is an n-dimensional vector of 1’s.

C. Fernández gratefully acknowledges financial support from a Training and Mobility of Researchers grant awarded by the European Commission (ERBFMBICT # 961021). C. Fernández and M.F.J. Steel were affiliated to CentER and the Department of Econometrics, Tilburg University, The Netherlands during the early stages of the work on this paper.

This gives rise to possible sampling models, depending on whether we include or exclude each of the regressors. A model , consists of a choice of regressors and leads to

\[y = \alpha \iota_ {n} + Z _ {j} \beta_ {j} + \sigma \varepsilon ,\tag{1.1}\]

where is the vector of observations. In (1.1), denotes the submatrix of of relevant regressors, groups the corresponding regression coefficients and is a scale parameter. Furthermore, we shall assume that ε follows an n-dimensional Normal distribution with zero mean and identity covariance matrix.

Wenow need tospecifyapriordistributionfortheparametersin , namely and . Thisdistribution will be given through a density function

\[p (\alpha , \beta_ {j}, \sigma \mid M _ {j}).\tag{1.2}\]

In Section 2, we shall consider specific choices for the density in (1.2) and examine the resulting Bayes factors. In order to complete the prior distribution of the parameters under model , we group the irrelevant components of under in a vector . The latter vector follows a Dirac distribution at zero, i.e.

\[P _ {\beta_ {\sim j} \mid \alpha , \beta_ {j}, \sigma , M _ {j}} = P _ {\beta_ {\sim j} \mid M _ {j}} = \mathrm{Diracat} (0, \ldots , 0).\tag{1.3}\]

We denote the space of all possible models by thus

\[\mathcal {M} = \{M _ {j}: j = 1, \dots , 2 ^ {k} \}.\tag{1.4}\]

In a Bayesian framework, dealing with model uncertainty is, theoretically, perfectly straightforward: we simply need to put a prior distribution over the model space M

\[P (M _ {j}) = p _ {j}, j = 1, \dots , 2 ^ {k}, \quad \text { with } p _ {j} > 0 \text { and } \sum_ {j = 1} ^ {2 ^ {k}} p _ {j} = 1.\tag{1.5}\]

The Bayesian model is then specified in three consecutive steps:

(1) Through (1.1) we define the distribution of the observables given model and the parameters and σ.

(2) In (1.2)-(1.3) we specify the distribution of the parameters and σ given

(3) Finally, (1.5) gives the prior probabilities of each of the models.

With this setup, the posterior distribution of any quantity of interest, say is a mixture of the posterior distributions of that quantity under each of the models with mixing probabilities given by the posterior model probabilities. Thus

\[P _ {\Delta \mid y} = \sum_ {j = 1} ^ {2 ^ {k}} P _ {\Delta \mid y, M _ {j}} P (M _ {j} \mid y).\tag{1.6}\]

This procedure, which is typically referred to as Bayesian model averaging (BMA), is in fact the standard Bayesian procedure under model uncertainty, since it follows directly from the rules ofprobability calculus upon which the Bayesian paradigm is based [see e.g. Leamer (1978), Min and Zellner (1993), Osiewalski and Steel (1993) and Raftery et al. (1997)].

Posterior model probabilities are given by

\[P (M _ {j} \mid y) = \frac {l _ {y} (M _ {j}) P (M _ {j})}{\sum_ {h = 1} ^ {2 ^ {k}} l _ {y} (M _ {h}) P (M _ {h})} = \left(\sum_ {h = 1} ^ {2 ^ {k}} \frac {P (M _ {h})}{P (M _ {j})} \frac {l _ {y} (M _ {h})}{l _ {y} (M _ {j})}\right) ^ {- 1},\tag{1.7}\]

where , the marginal likelihood ofmodel , is obtained as

\[l _ {y} (M _ {j}) = \int p (y \mid \alpha , \beta_ {j}, \sigma , M _ {j}) p (\alpha , \beta_ {j}, \sigma \mid M _ {j}) d \alpha d \beta_ {j} d \sigma ,\tag{1.8}\]

with and defined through (1.1) and (1.2), respectively.

Two difficult questions here are how to compute and how to assess the influence ofour prior assumptions on the latter quantity. In cases where can be derived analytically, the computation of is reasonably straightforward applying the methodology of Madigan and York (1995). This is a Metropolis algorithm [see Chib and Greenberg (1995)], which allows us to generate drawings from a Markov chain on the model space M with the posterior model distribution as its stationary distribution. The latter is easily implemented if combined with the choice of a natural conjugate prior structure [see e.g. Raftery et al. (1997)]. For more complex prior structures that do not allow for an explicit expression for , the reversiblejump methodology of Green (1995) could be applied. An alternative approach was proposed by George and McCulloch (1993, 1997), who do not formally impose zero restrictions as in (1.3), but constrain these coefficients to be concentrated around zero instead (in this way, they get around the problem of a parameter space of varying dimension).

Ontheotherhand, theissueofchoosinga“sensible” priordistributionseemsfurtherfrombeingresolved. From (1.7) it is clear that the value of is determined by the prior odds and the Bayes factors ofeach ofthe entertained models versus . Bayes factors are known to be rather sensitive to the choice of the prior distributions for the parameters within each model. Even asymptotically, theinfluenceofthis distributiondoes notvanish[see Kassand Raftery (1995) and George (1997)]. Thus, under little (or under absence of) prior information, the choice ofthe distribution in (1.2) is a very thorny question. Furthermore, the usual recourse to improper “noninformative” priors does not work in this situation, since the rules of probability no longer apply if we use improper priors on model-specific parameters. Most of the priors that have been proposed in the literature violate the rules of probability calculus, since they are either improper on model-specific parameters [attempts to overcome this are intrinsic Bayes factors as in Berger and Pericchi (1996) or fractional Bayes factors as in O’Hagan (1995)] or are data-dependent through the response variable [as the prior in Raftery et al. (1997)]. Here, we will focus on priors that do not have these undesirable properties and are thus, in our view, more suitable for a Bayesian analysis. To this end, we shall propose certain priors and study their behaviour in comparison with other priors previously considered in the literature in Bernardo (1980) and Laud and Ibrahim (1995, 1996)].

2. PRIORS FOR MODEL PARAMETERS AND THE CORRESPONDING BAYES FACTORS

In this section, we present several priors [i.e. several choices for the density in (1.2)] and derive the expressions of the resulting Bayes factors. In the next sections, we examine the properties (both finitesample and asymptotic) of the Bayes factors.

2.1. A natural conjugate framework

Both for reasons of computational simplicity and for the interpretability of theoretical results, the most obvious choice for the prior distribution of the parameters is a natural conjugate one. The density in (1.2) is then given through

\[p (\alpha , \beta_ {j} \mid \sigma , M _ {j}) = f _ {N} ^ {k _ {j} + 1} ((\alpha , \beta_ {j}) \mid m _ {0 j}, \sigma^ {2} V _ {0 j}),\tag{2.1}\]

which denotes the p.d.f. of a )-variate Normal distribution with mean and covariance matrix , and through

\[p (\sigma^ {- 2} \mid M _ {j}) = p (\sigma^ {- 2}) = f _ {G} (\sigma^ {- 2} \mid c _ {0}, d _ {0}),\tag{2.2}\]

which corresponds to aGammadistribution with mean and variance for . Notethatwehave assumed a common prior distribution for σ across models. Clearly PDS matrix, and are prior hyperparameters that still need to be elicited.

This natural conjugate framework greatly facilitates the computation of posterior distributions and Bayes factors. In particular, the marginal likelihood of model computed through (1.8) takes the form

\[l _ {y} (M _ {j}) = f _ {S} ^ {n} \left(y \mid 2 c _ {0}, X _ {j} m _ {0 j}, \frac {c _ {0}}{d _ {0}} \left(I _ {n} - X _ {j} V _ {* j} X _ {j} ^ {\prime}\right)\right),\tag{2.3}\]

where

\[X _ {j} = (\iota_ {n}: Z _ {j}),\tag{2.4}\]

\[V _ {* j} = (X _ {j} ^ {\prime} X _ {j} + V _ {0 j} ^ {- 1}) ^ {- 1},\tag{2.5}\]

and denotes the p.d.f. of an n-variate Student-t distribution with ν degrees of freedom, location vector b (the mean if ) and precision matrix A (with covariance matrix provided evaluated at y. The Bayes factor for model versus model now takes the form

\[B _ {j s} = \frac {l _ {y} (M _ {j})}{l _ {y} (M _ {s})} = \left(\frac {| V _ {* j} |}{| V _ {0 j} |} \frac {| V _ {0 s} |}{| V _ {* s} |}\right) ^ {1 / 2} \left\{\frac {2 d _ {0} + (y - X _ {s} m _ {0 s}) ^ {\prime} (I _ {n} - X _ {s} V _ {* s} X _ {s} ^ {\prime}) (y - X _ {s} m _ {0 s})}{2 d _ {0} + (y - X _ {j} m _ {0 j}) ^ {\prime} (I _ {n} - X _ {j} V _ {* j} X _ {j} ^ {\prime}) (y - X _ {j} m _ {0 j})} \right\} ^ {c _ {0} + \frac {n}{2}}.\tag{2.6}\]

Generally, the choice of the prior hyperparameters in (2.1)-(2.2) is not a trivial one. The user is plagued by the pitfalls described in Richard (1973), arising if we wish to combine a fixed quantity of subjective prior information on the regression coefficients with little prior information on σ. Richard and Steel (1988, App. D) and Bauwens (1991) propose a subjective elicitation procedure for the precision parameter based on the expected fit of the model. See Poirier (1996) for related ideas. In this paper we shall follow the opposite strategy, and instead of trying to elicit more prior information in a situation of incomplete prior specification, we focus on situations where we have (or wish to use) as little subjective prior knowledge as possible.

2.2. Choosing prior hyperparameters for

Choosing and can be quite difficult in the absence of prior information. A predictive way of eliciting is through making a prior guess for the n-dimensional response y. Laud and Ibrahim (1996) propose to make such a guess, call it taking the information on all the covariates into account and subsequently choose . Our approach is similar in spirit but much simpler: Given that we do not possess a lot of prior information, we consider it very difficult to make a prior guess for n observations taking the covariates for each ofthese n observations into account. Especially when n is large, this seems like an extremely demanding task. Instead, one could hope to have an idea ofthe central values of and make the following prior prediction guess: , which corresponds to

\[m _ {0 j} = (m _ {1}, 0, \dots , 0) ^ {\prime}.\tag{2.7}\]

Eliciting prior correlations is even more difficult. We adopt the popular and convenient g-prior [Zellner (1986)], which corresponds to taking

\[V _ {0 j} ^ {- 1} = g _ {0 j} X _ {j} ^ {\prime} X _ {j},\tag{2.8}\]

with . From (2.5) it is clear that is the prior counterpart of . This choice is extremely popular, and has been considered, among others by Poirier (1985) and Laud and Ibrahim (1995, 1996). See also Smith and Spiegelhalter (1980) for a closely related idea.

With these hyperparameter choices, the Bayes factor in (2.6) can be written in the following intuitively interpretable way

\[B _ {j s} = \left(\frac {g _ {0 j}}{g _ {0 j} + 1}\right) ^ {\frac {k _ {j} + 1}{2}} \left(\frac {g _ {0 s} + 1}{g _ {0 s}}\right) ^ {\frac {k _ {s} + 1}{2}} \left(\frac {2 d _ {0} + \frac {1}{g _ {0 s} + 1} y ^ {\prime} M _ {X _ {s}} y + \frac {g _ {0 s}}{g _ {0 s} + 1} (y - m _ {1} \iota_ {n}) ^ {\prime} (y - m _ {1} \iota_ {n})}{2 d _ {0} + \frac {1}{g _ {0 j} + 1} y ^ {\prime} M _ {X _ {j}} y + \frac {g _ {0 j}}{g _ {0 j} + 1} (y - m _ {1} \iota_ {n}) ^ {\prime} (y - m _ {1} \iota_ {n})}\right) ^ {c _ {0} + \frac {n}{2}},\tag{2.9}\]

where

\[y ^ {\prime} M _ {X _ {j}} y = y ^ {\prime} y - y ^ {\prime} X _ {j} (X _ {j} ^ {\prime} X _ {j}) ^ {- 1} X _ {j} ^ {\prime} y\tag{2.10}\]

is the usual Sum ofSquared Residuals under model

Note that the last factor in (2.9) contains a convex combination between the model “lack offit” (measured through and the “error ofour prior prediction guess” [measured through The coefficients ofthis convex combination are determined by the choice of . The choice of is crucial for obtaining sensible results, as we shall see later. By not choosing through fixing a marginal prior of the regression coefficients, we avoid the natural conjugate pitfall mentioned at the end ofSubsection 2.1. In addition, the -prior in (2.8) can also lead to a prior that is continuously induced across models [see Poirier (1985)] in the sense that the priors for all J models can be derived as the relevant conditionals from the prior of the full model (with . This will hold as long as does not depend on

2.3. A non-informative prior for σ

From (2.9) itis clear thatthechoiceof , theprecisionparameterintheGammapriordistributionfor can crucially affect the Bayes factor. In particular, ifthe value of is large in relation to the values of and the prior will dominate the sample information, which is a rather undesirable property. The impact of on the Bayes factor also clearly depends on the units of measurement for the data y. In the absence of(or under little) prior information, it is very difficult to choose this hyperparameter value without using the data ifwe do not want to risk choosing it too large. Even using prior ideas about fit does not help; Poirier (1996) shows that the population analog ofthe coefficient ofdetermination does not have any prior dependence on or . Use of the information in the response variable was proposed by Raftery (1996) and Raftery et al. (1997) but, as we already mentioned, this takes us outside the rules of probability and we prefer to avoid this situation. Instead we propose the following:

Since the scale parameter σ appears in all the models entertained, we can use the improper prior distribution with density

\[p (\sigma) \propto \sigma^ {- 1},\tag{2.11}\]

which is the widely accepted non-informative prior distribution for scale parameters. It is easy to check that this improper prior leads to a proper posterior (and thus allows for a Bayesian analysis) as long as . The distribution in (2.11) is the only one that is invariant under scale transformations (induced by a change in the units of measurement) and is the limiting distribution of the Gamma conjugate prior in (2.2) when both and tend to zero. This leads to the Bayes factor

\[B _ {j s} = \left(\frac {g _ {0 j}}{g _ {0 j} + 1}\right) ^ {\frac {k _ {j} + 1}{2}} \left(\frac {g _ {0 s} + 1}{g _ {0 s}}\right) ^ {\frac {k _ {s} + 1}{2}} \left(\frac {\frac {1}{g _ {0 s} + 1} y ^ {\prime} M _ {X _ {s}} y + \frac {g _ {0 s}}{g _ {0 s} + 1} (y - m _ {1} \iota_ {n}) ^ {\prime} (y - m _ {1} \iota_ {n})}{\frac {1}{g _ {0 j} + 1} y ^ {\prime} M _ {X _ {j}} y + \frac {g _ {0 j}}{g _ {0 j} + 1} (y - m _ {1} \iota_ {n}) ^ {\prime} (y - m _ {1} \iota_ {n})}\right) ^ {\frac {n}{2}},\tag{2.12}\]

where we have avoided the influence ofthe hyperparameter values and .

2.4. A non-informative prior for the intercept

In (2.12) there are two subjective elements that still remain, namely the choices of and of , where is our prior guess for . It is clear from (2.12) that the choice of can have a non-negligible impact on the actual Bayes factor and, under absence of prior information, it is extremely difficult to successfully elicit without using the data. The idea that we propose here is very much in line with our solution for the prior on since all the models have an intercept, take the usual non-informative improper prior for a location parameter with constant density. This avoids the difficult issue of choosing a value for

This setup takes us outside the natural conjugate framework, since our prior for no longer corresponds to (2.1). Without loss of generality, we assume that

\[\iota_ {n} ^ {\prime} Z = 0,\tag{2.13}\]

so that the intercept is orthogonal to all the regressors. This is immediately achieved by substracting the corresponding meanfromeachofthem. Suchatransformationonly affectstheinterpretationoftheintercept which is typically not of primary interest. In addition, the prior that we next propose for α is not affected by this transformation. We now consider the following prior density for :

\[p (\alpha) \propto 1,\tag{2.14}\]

\[p (\beta_ {j} \mid \sigma , M _ {j}) = f _ {N} ^ {k _ {j}} (\beta_ {j} \mid 0, \sigma^ {2} (g _ {0 j} Z _ {j} ^ {\prime} Z _ {j}) ^ {- 1}).\tag{2.15}\]

Through we assume the same prior distribution for α in all of the models and a g-prior distribution for under model . We again use the non-informative prior described in for Existence of a proper posterior distribution is now achieved as long as the sample contains at least two different observations. The Bayes factor for versus now is

\[B _ {j s} = \left(\frac {g _ {0 j}}{g _ {0 j} + 1}\right) ^ {k _ {j} / 2} \left(\frac {g _ {0 s} + 1}{g _ {0 s}}\right) ^ {k _ {s} / 2} \left(\frac {\frac {1}{g _ {0 s} + 1} y ^ {\prime} M _ {X _ {s}} y + \frac {g _ {0 s}}{g _ {0 s} + 1} (y - \overline {{y}} \iota_ {n}) ^ {\prime} (y - \overline {{y}} \iota_ {n})}{\frac {1}{g _ {0 j} + 1} y ^ {\prime} M _ {X _ {j}} y + \frac {g _ {0 j}}{g _ {0 j} + 1} (y - \overline {{y}} \iota_ {n}) ^ {\prime} (y - \overline {{y}} \iota_ {n})}\right) ^ {(n - 1) / 2},\tag{2.16}\]

if and . Ifone ofthe latter two quantities, , is zero (which corresponds to the model with just the intercept), the Bayes factor is simply obtained as the limit of in (2.16) letting tend to infinity.

Note the similarity between the expression in (2.16) and (2.12), where we had adopted a (limiting) natural conjugate framework. When we are non-informative on the intercept [see (2.16)] we lose, as it were, one observation becomes and one regressor becomes . But the most important difference is that our subjective prior guess is now replaced by which is eminently reasonable and avoids the sensitivity problems alluded to before. Thus, we shall, henceforth, favour the prior given by the product of (2.11), (2.14) and (2.15).

3. ASYMPTOTIC PROPERTIES AND THE CHOICE OF

Following our comment at the end ofSection 2, the remainder ofthe paper will focus on the prior given through (2.11), (2.14) and (2.15), which leads to the expression in (2.16) for the Bayes factor. We note that in (2.16) only remains to be determined. This will be done using a number of properties of the Bayes factor and the posterior model probabilities. In particular, we would like to have consistency (i.e. assuming that one of the entertained models is the correct one, we would want the posterior probability of the correct model to converge to one as sample size increases). In addition, we also want sensible behaviour for finite sample sizes, both in terms of posterior model probabilities and predictive ability. In this section we shall focus on large sample results, whereas Section 4 will deal with finite-sample properties through a simulation experiment.

Throughout this section we assume that the sample is generated by model with parameter values and

\[y = \alpha \iota_ {n} + Z _ {s} \beta_ {s} + \sigma \varepsilon .\tag{3.1}\]

We aim at achieving consistency in the sense that

\[\underset {n \to \infty} {\text {plim}} P (M _ {s} \mid y) = 1 \text {and} \underset {n \to \infty} {\text {plim}} P (M _ {j} \mid y) = 0 \text {for all} M _ {j} \neq M _ {s},\tag{3.2}\]

where the probability limit is taken with respect to the true sampling distribution described in (3.1). By (1.7), as long as the prior (1.5) on the model space does not depend on sample size, we simply need to check that the Bayes factor for model versus model , converges in probability to zero for any model other than . The reference posterior odds proposed in Bernardo (1980) and Pericchi (1984) rely on making prior model probabilities depend on the expected gain in information from the sample. As explained in these papers, such procedures will generally not lead to consistency in the sense of (3.2).

Although we shall focus on the case ofimproper priors on α and σ, thus leading to the expression for in (2.16), it is immediate to see that the same results apply to the Bayes factor in (2.9) (which corresponds to proper priors on both and and to the Bayes factor in (2.12) (where we are still proper on α).

The appendix will group some derivations underlying the results in this section. We shall assume throughout that condition (A.2) in the appendix holds. We examine two different functional choices for Let us first consider dependence on the sample size n and, possibly, on the number of regressors

\[\text { 3.1. Results under } g _ {0 j} = \frac {w _ {1} (k _ {j})}{w _ {2} (n)} \text { with } \lim _ {n \to \infty} w _ {2} (n) = \infty\]

This is a rather logical choice for in view ofthe prior in (2.15), which assumes that the prior precision is a fraction of the sample precision. Thus, it seems natural to impose that as sample size increases, the precision ofthe prior becomes a smaller fraction ofthat ofthe sample and vanishes as n goes to infinity. In addition, we let depend on a function of . The following theorem summarizes our results:

Theorem 1. Consider the Bayesian model given by (1.1), together with the prior densities in (2.11), (2.14), (2.15) and any prior on the model space in (1.5). We assume that in (2.15) takes the form

\[g _ {0 j} = \frac {w _ {1} (k _ {j})}{w _ {2} (n)} w i t h \lim _ {n \to \infty} w _ {2} (n) = \infty .\tag{3.3}\]

Then, under the assumption that there is a true model in M that generates the data, the condition

\[\lim _ {n \to \infty} \frac {w _ {2} ^ {\prime} (n)}{w _ {2} (n)} = 0,\tag{3.4}\]

together with either

\[\lim _ {n \to \infty} \frac {n}{w _ {2} (n)} \in [ 0, \infty)\tag{3.5}\]

or

\[w _ {1} (\cdot) \text { is a nondecreasing function, }\tag{3.6}\]

ensures that the posterior distribution of the models is consistent in the sense defined in (3.2).

On the basis of prior ideas about fit, Poirier (1996) suggests taking , which satisfies (3.4) and (3.5) and thus leads to consistent Bayes factors. This and other choices will be discussed subsequently.

The proof of Theorem 1 (see appendix) never makes use of the Normality assumption for the error distribution of the ‘true’ model in (3.1), and, thus, our findings immediately generalize to the case where the components ofε in (3.1) are i.i.d. following any regular distribution with finite variance. Therefore, even ifthe true model does not possess a Normally distributed error term, the posterior distribution derived on the basis of the models with Normality assumed [leading to the Bayes factor in (2.16)] is still consistent, in the sense ofasymptotically selecting the true subset ofregressors, under the sufficient conditions for stated in Theorem 1. This implies that we can always make the convenient assumption of Normality to asymptotically select the correct set of regressors. In some sense, this offers a counterpart to the classical resultfortestingnested models, wheretheLikelihood Ratio, Wald and Rao(orLagrangemultiplier) statistics derived under the assumption of Normality keep the same asymptotic distribution (a even if the error term is non-Normal [see Amemiya (1985, p. 144)].

\[3. 2. \text { Results with } g _ {0 j} = w (k _ {j})\]

We now examine the situation where is no longer a function of the sample size n. Therefore, consistency is entirely driven by the last factor of in (2.16), which we denote by .

It is immediately clear that in this situation we do not have consistency: When the data generating model, , is the model withjust the intercept, regardless of the data [since the numerator in the last factor of is then , which is always bigger than or equal to the denominator]. Thus, can not converge to one as n tends to infinity, precluding consistency.

Even though we do not have consistency, let us examine the asymptotic behaviour of for the case where contains some regressors other than the intercept (i.e. . See the appendix for proofs.

When is nested within , we have the following result:

\[\underset {n \to \infty} {\mathrm{plim}} D _ {j s} = 0 \text {if and only if} w (\cdot) \text {is an increasing function.}\tag{3.7}\]

The situation becomes less clear-cut when is not nested within . We can show that converges to zero when is the model withjust the intercept. In addition, taking to be an increasing function, is sufficient if , but we can not assure this if (this will be case-specific). Thus, we can not exclude that models smaller than the true one asymptotically receive positive posterior probability.

On the other hand, if we take to be a constant, we obtain a zero limit for if is not nested within . However, in this situation models that nest the true model asymptotically receive positive probability, as follows from (3.7).

3.3. Relationship to information criteria

A number of information criteria have traditionally been used for classical model selection purposes, especially in the area of time series analysis. In this subsection, we shall establish asymptotic links between the Bayes factors corresponding to Subsection 3.1 and two consistent information criteria: the Schwarz (or Bayes information) criterion as derived in Schwarz (1978) and the Hannan-Quinn criterion ofHannan and Quinn (1979). Ifwe wish to compare two models as in (1.1), say versus , these criteria take the form:

\[S _ {j s} = \frac {n}{2} \ln \left(\frac {y ^ {\prime} M _ {X _ {s}} y}{y ^ {\prime} M _ {X _ {j}} y}\right) + \frac {k _ {s} - k _ {j}}{2} \ln (n),\tag{3.8}\]

\[H Q _ {j s} = \frac {n}{2} \ln \left(\frac {y ^ {\prime} M _ {X _ {s}} y}{y ^ {\prime} M _ {X _ {j}} y}\right) + \frac {k _ {s} - k _ {j}}{2} C _ {H Q} \ln \ln (n).\tag{3.9}\]

Hannan and Quinn (1979) prove strong consistency for both criteria provided

The asymptotic behaviour ofthe Bayes factor in (2.16), made consistent by choosing as in Theorem 1, can be characterized by the following result:

Theorem 2. Consider the Bayesian model described in Theorem 1, with verifying (3.4) together with either (3.5) or (3.6). Then the Bayes factor in (2.16) satisfies:

\[\mathrm{plim} \frac {\ln B _ {j s}}{\frac {n}{2} \ln \left(\frac {y ^ {\prime} M _ {X _ {s}} y}{y ^ {\prime} M _ {X _ {j}} y}\right) + \frac {k _ {s} - k _ {j}}{2} \ln w _ {2} (n)} = 1,\tag{3.10}\]

where the probability limit is taken with respect to the model as described in (3.1).

Thus, different choices ofthe function will influence the asymptotic behaviour ofthe logarithm of the Bayes factor. In particular, let us consider the choices of that induce a relationship with the two information criteria mentioned above.

Corollary 1. If in Theorem 2 we choose , we obtain

\[\mathrm{plim} \frac {\ln B _ {j s}}{S _ {j s}} = 1,\tag{3.11}\]

whereas choosing and nondecreasing, leads to

\[\mathrm{plim} \frac {\ln B _ {j s}}{H Q _ {j s}} = 1.\tag{3.12}\]

Fromtheseresultsweseethatln behavesliketheseconsistentcriteriaifwechoose appropriately. Note that the second choice of in Corollary 1 does not verify (3.5), which is why we impose that fulfills (3.6). Kass and Wasserman (1995) study the relationship between the Schwarz criterion and Bayes factors using “unit information priors” for testing nested hypotheses, and provide the order of the approximation under certain regularity conditions.

As a final note, it is again worth mentioning that Theorem 2also holds ifthe error terms in (3.1) follow a non-Normal distribution.

4. THE SIMULATION EXPERIMENT

4.1. Introduction

In this section we perform a simulation experiment to assess the performance of different choices of in finite sampling. In addition to Bayes factors, we will compute posterior model probabilities and evaluate predictive ability under several choices of . Our results in this section will be derived under a Uniform prior on the model space M. Thus, the Bayesian model will be given through (1.1), together with the prior densities in (2.11), (2.14) and (2.15), and

\[P (M _ {j}) = p _ {j} = 2 ^ {- k}, \quad j = 1, \ldots , k.\tag{4.1}\]

Creating the design matrix ofthe simulation experiment follows Example 5.2.2in Raftery et al. (1997). We generate an matrix R ofregressors in the following way: the first ten columns in R, denoted are drawn from independent standard Normal distributions, and the next five columns are constructed from

\[(r _ {(1 1)}, \dots , r _ {(1 5)}) = (r _ {(1)}, \dots , r _ {(5)}) (. 3 \quad . 5 \quad . 7 \quad . 9 \quad 1. 1) ^ {\prime} (1 \quad 1 \quad 1 \quad 1 \quad 1) + E\tag{4.2}\]

where E is an matrix ofindependent standard Normal deviates. Note that (4.2) induces a correlation between the first five regressors and the last five regressors. The latter takes the form of small to moderate correlations between , and (the theoretical correlation coefficients increase from 0.153 to 0.561 with i) and somewhat larger correlations between the last five regressors (theoretical values 0.740). Curiously, this correlation structure differs from the one reported in Raftery et al. (1997), which seems in conflict with (4.2). After generating R, we demean each ofthe regressors, thus leading to a matrix that fulfills (2.13).

A vector of n observations is then generated according to one of the models

\[\text { Model 1: } y = 4 + 2 z _ {(1)} - z _ {(5)} + 1. 5 z _ {(7)} + z _ {(1 1)} + 0. 5 z _ {(1 3)} + u,\tag{4.3}\]

\[\text { Model 2: } \quad y = 1 + u,\tag{4.4}\]

where the elements of u are independently Normally distributed with mean zero and variance Whereas Model 1 is meant to capture a more or less realistic situation where one third of the regressors intervene, Model 2 is an extreme case without any relationship between predictors and response. A “null model” similar to the latter was analysed in Freedman (1983) using a classical approach and in Raftery et al. (1997) through Bayesian model averaging.

4.2. Choices for

Based on the theory in Section we shall consider seven different choices for . From Theorem 1, priors a-f all lead to consistency, in the sense of asymptotically selecting the correct model. In addition, from Corollary 1, the log ofBayes factors obtained under priors a-c behave asymptotically like the Schwarz criterion, whereas those obtained under priors e and frespectively behave like the Hannan-Quinn criterion with and . Prior d provides an intermediate case in terms of asymptotic penalty for large models.

Prior a:

This prior roughly corresponds to assigning the same amount of information to the conditional prior of as is contained in one observation. Thus, it is in the spirit of the “unit information priors” of Kass and Wasserman (1995) and the g-prior (using a Cauchy prior on given used in Zellner and Siow (1980). In addition, this choice of does not depend on j and, thus, implies continuously induced priors across models, in the sense of Poirier (1985).

Prior b:

Here we assign more information to the prior as we have more regressors in the model, i.e. we expect the sample information to be more diluted as the number of regressors grows.

Prior c:

Now prior information decreases with the number of regressors in the model.

Prior d:

This is an intermediate case, where we choose , and we have a smaller asymptotic penalty term for large models than in the Schwarz criterion [see (3.8)].

Prior e:

Here we choose so as to mimic the Hannan-Quinn criterion in (3.9) with as n becomes large. This prior is also continuously induced across models.

Prior f:

Now increases even slower with sample size and we have asymptotic convergence ofln to s for

Prior g:

This choice was suggested by Laud and Ibrahim (1996), who use a natural conjugate prior structure, subjectively elicited through predictive implications. In applications, they propose to choose and δ such that (the weight of the “prior prediction error” in our Bayes factors); for this implies: . Note that this prior choice does not lead to consistency if the data are generated from a model with only the intercept, the “null model” in (4.4) (see Subsection 3.2). If then is an increasing function of , and following Subsection 3.2we know that plim when and is either the null model or . In cases when it depends on (A.9) whether converges to zero or not.

4.3. Predictive criteria

Clearly, ifwe generate the data from some known model, we are interested in recovering that model with the highest possible posterior probability for each given sample size . However, in practical situations with real data, we might be more interested in predicting the observable, rather than uncovering some “true” underlying structure. This is more in line with the Bayesian way ofthinking, where models are mere “windows” through which to view the world [see Poirier (1988)], but have no inherent meaning in terms of characteristics of the real world. See also Dawid (1984) and Geisser and Eddy (1979).

Forecasting is conducted conditionally upon the regressors, so we will generate k-dimensional vectors , given which we will predict the observable In empirical applications, will typically be constructed from some original value of which we substract the mean of the raw regressors in the sample on which inference is based. This ensures that the interpretation of the regression coefficients in posterior and predictive inference is compatible.

In this subsection, it will prove useful to explicit the conditioning on the regressors in and in the notation. In accordance with the expression in (1.6), the out-of-sample predictive distribution for will be characterized by

\[\begin{array}{r l} & p (y _ {f} \mid z _ {f}, y, Z) = \sum_ {j = 1} ^ {J} f _ {S} ^ {1} (y _ {f} \mid n - 1, \overline {{y}} + \frac {1}{g _ {0 j} + 1} z _ {f, j} ^ {\prime} \beta_ {j} ^ {*}, \\ & \qquad \frac {n - 1}{d _ {j} ^ {*}} \{1 + \frac {1}{n} + \frac {1}{g _ {0 j} + 1} z _ {f, j} ^ {\prime} (Z _ {j} ^ {\prime} Z _ {j}) ^ {- 1} z _ {f, j} \} ^ {- 1}) P (M _ {j} \mid y, Z), \end{array}\tag{4.5}\]

where is based on the inference sample groups the elements of corresponding to the regressors in and

\[d _ {j} ^ {*} = \frac {1}{g _ {0 j} + 1} y ^ {\prime} M _ {X _ {j}} y + \frac {g _ {0 j}}{g _ {0 j} + 1} (y - \overline {{y}} \iota_ {n}) ^ {\prime} (y - \overline {{y}} \iota_ {n})\tag{4.6}\]

The term in (4.5) corresponding to the model with only the intercept is obtained by letting the corresponding tend to infinity.

The log predictive score is a proper scoring rule introduced by Good (1952). Some of its properties are discussed in Dawid (1986). This is the first predictive criterion we will compute. For each value of we shall generate a number, say of responses from the underlying true model [(4.3) or (4.4)] and base our predictive measure on (4.5) evaluated in these out-of-sample observations , namely:

\[L P S (z _ {f}, y, Z) = - \frac {1}{v} \sum_ {i = 1} ^ {v} \ln p (y _ {f i} \mid z _ {f}, y, Z),\tag{4.7}\]

It is clear that a smaller value of makes a Bayes model (thus, in our context, a prior choice for preferable. Madigan, Gavrin and Raftery (1995) give an interpretation for differences in log predictive scores in terms of one toss with a biased coin.

More formally, the criterion in (4.7) can be interpreted as an approximation to the expected loss with a logarithmic rule, which is linked to the well-known Kullback-Leibler criterion. The Kullback-Leibler divergence between the actual sampling density in (4.3) or (4.4) and the out-of-sample predictive density in (4.5) can be written as

\[\begin{array}{c} K L \{p (y _ {f} \mid z _ {f}), p (y _ {f} \mid z _ {f}, y, Z) \} = \int_ {\Re} \{\ln p (y _ {f} \mid z _ {f}) \} p (y _ {f} \mid z _ {f}) d y _ {f} - \\ \int_ {\Re} \{\ln p (y _ {f} \mid z _ {f}, y, Z) \} p (y _ {f} \mid z _ {f}) d y _ {f}, \end{array}\tag{4.8}\]

where the first integral is the negative entropy of the sampling density, and the second integral can be seen as a theoretical counterpart of (4.7) for a given value of . This latter integral can easily be shown to be finite in our particular context and is now approximated by averaging over v values for given a particular vector of regressors . For the Normal sampling model used here, the negative entropy is given by .335 for our choice of in (4.3), regardless of . By the nonnegativity of the Kullback-Leibler divergence, this constitutes a lower bound for of2.335.

We now have a measure of predictive performance, , for each . We can, of course, immediately assess the distribution of this quantity as varies, through its empirical distribution corresponding to all generated values of

Finally, we can investigate the calibration of the predictive and compare the entire predictive density function in (4.5) with the known sampling distribution of the response in (4.3) or (4.4) given a particular (fixed) set of regressor variables. The fact that such predictions are, by the very nature of our regression model, conditional upon the regressors does complicate matters slightly. We can not simply compare the sampling density averaged over different values of with the averaged predictive density function. It is clearly crucial to identify predictives with the value of they condition on. Predicting correctly “on average” can mask arbitrarily large errors in conditional predictions, as long as they compensate each other. We shall graphically present comparisons of the sampling density and the predictive density for three key values within our sample of predictors: the one leading to the smallest mean ofthe sampling model in (4.3), the one leading to the median value and the one giving rise to the largest value. For the sampling model in (4.4), the value of does not intervene, so here we shalljust present a graph ofthe predictive for one (randomly chosen) value (the sampling model now does not change with , but the predictive does, as long as we assign nonzero probability to models larger than the null model). In addition, we shall present properties of and the predictive coverage averaged over the different values of as well. These latter measures of predictive performance naturally compare each predictive with the corresponding sampling distribution (i.e. taking the value of into account), so that an overall measure can readily be computed.

5. FINITE SAMPLE SIMULATION RESULTS

5.1 Convergence and implementation

The implementation ofthe simulation study described in the previous section will be conducted through the methodology mentioned in Section 1. This Metropolis algorithm generates a new candidate model, say , from a Uniform distribution over the subset of M consisting of the current state of the chain, say and all models containing one regressor more or less than . The chain moves to with probability min , where is the Bayes factor in (2.16). In order to evaluate the posterior model probabilities we can simply count the relative frequencies of model visits in the Markov chain. A somewhat more interesting alternative to this strategy is to use the actual Bayes factors, already computed in running the chain, to compare all visited models. Since the number of visited models is typically a small subset ofthe total number ofpossible models, this method is feasible. Lee (1996) introduces this idea as Bayesian Random Search (BARS). The generated chain is then effectively only used to indicate which models should be considered in computing Bayes factors. All other (non-visited) models will implicitly be assumed to have zero posterior probability. This has two advantages: firstly, it is clearly more precise than relative frequencies, since the Bayes factors in (2.16) are exact and don’t require any ergodic properties. Secondly, comparing empirical relative frequencies with exact Bayes factors will give a good indication of the convergence of the chain. We shall report results based on Bayes factors, but we ran the chain for long enough to get almost the same answers with empirical model frequencies. This resulted in Markov Chain MonteCarlowith50,000recorded drawingsafteraburn-inof20,000drawings. A useful diagnostictoassess convergence ofthe Markov chain is the correlation coefficient ofthe exact Bayes factors computed through (2.16) and the relative frequencies of model visits.

In order to avoid results depending on the particular sample analyzed, we have generated 100independent samples according to the setup described in Section 4. Frequently, results will be presented in the form of either means and standard deviations or quantiles computed over these 100 samples. Sample sizes used in the simulation will be , 100, 1000, 10,000and 100,000. Furthermore, we generate different vectors of regressors for the forecasts of Model 1, whereas for Model 2. For each of these values of the vector out-of-sample observations will be generated.

As such a simulation study is quite CPU demanding, we put a good deal of emphasis on efficient coding and speed of execution. We coded in standard Fortran , and we used stacks to store information pertaining evaluated models in order to reduce the number of calculations. The entire simulation was run on one single 120MHz 604 PowerPC-based desktop computer. On a PowerMacintosh 7600, each 20,000– 50,000chain would take an average (over priors) time in seconds of: 209, 58, 5, 18, and 117; for 1000, 10,000and 100,000. The source code is posted at http://econwpa.wustl.edu and it is freely available.

5.2 Posterior model inference

Section 3 contains some results concerning the asymptotic behaviour of our Bayesian analysis as the number of observations n goes to infinity. Here we summarize the main results for various finite values of n.

5.2.1. Results under Model 1

One of the main indicators of the performance of the Bayesian methodology is the posterior probability assigned to the model that has generated the data. Ideally, one would want this probability to be very high for small or moderate values of that are likely to occur in practice. Table 1 presents the means and standard deviations across the100samples of fortheposteriorprobability ofthetruemodel (Model 1).

Columns correspond to the five sample sizes used and rows order the different priors a through g introduced in Subsection 4.2. In order to puttheseresults in abetter perspective, notethattheprior model probability of each ofthe 215 possible models is equal and amounts to 3.052·10−5. We know from the theoretical results in Subsection 3.1 thatpriors a-fareconsistentin thesenseof(3.2). From Subsection 3.2, weremain inconclusive about consistency under prior g, since we cannot exclude that models with less regressors than the true one receive, asymptotically, positive posterior probability. However, our simulation results will suggest that consistency holds in our particular example. It is clear from Table 1 that the posterior probability of Model 1 varies greatly in finite samples. Whereas prior d already performs very well for n = 1000, getting average probabilities of the correct model upwards of 0.97, prior e only obtains a probability of 0.60 with a sample as large as 100,000. Apart from the absolute probability of the correct model, it is also important to examine how much posterior weight is assigned to Model 1 relative to other models. Therefore, Table 2 presents quartiles of the ratio between the posterior probability of the correct model and the highest posterior probability of any other model. It is clear that in most cases this ratio tends to be far above unity, which is reassuring as it tells us that the most favoured model will still be the correct one, even though it may not have a lot of posterior mass attached to it. For example, with n = 50 prior f only leads to a mean posterior probability of Model 1 of 0.002 but still favours the correct model to the next best. In fact, the correct model is always favoured in at least 75 of the 100 samples, even for small sample sizes. Note that this compares favourably to results in George and McCulloch (1993).

Table 1. Model 1: Means and Stds of the posterior probability of the true model.

n50100100010,000100,000
PriorMeanStdMeanStdMeanStdMeanStdMeanStd
a0.01280.01970.05750.06180.52930.14010.81110.09280.92540.0760
b0.00660.00910.03320.03380.44070.13730.76010.10640.90480.0841
c0.01100.01590.05190.05330.48600.13740.78530.09990.91450.0804
d0.00290.00260.02050.01880.97300.01961.00000.00001.00000.0000
e0.01410.02230.05860.06160.36100.12510.51390.13270.59810.1327
f0.00200.00140.01280.01070.77620.34211.00000.00001.00000.0000
g0.00260.00260.00690.00560.27730.08641.00000.00001.00000.0000

Table 2. Model 1: Quartiles of ratio of posterior probabilities; True Model vs Best among the rest.

$\mathcal{N}$ 50100100010,000100,000
PriorQ1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3
a1.53.26.33.05.88.69.119.029.827.667.890.579.0193.4285.5
b1.12.64.22.24.16.29.616.622.116.636.366.359.7113.4194.7
c1.23.45.82.05.18.37.013.822.926.355.973.662.3135.6236.1
d1.62.73.51.93.85.8226.5416.4629.4
e1.74.07.02.04.48.74.79.216.59.719.724.912.722.234.2
f1.42.33.31.43.95.111.4238.73625.8
g1.22.32.81.92.83.95.910.612.9

Table 3. Model 1: Means and Stds ofNumber ofModels Visited.

n50100100010,000100,000
PriorMeanStdMeanStdMeanStdMeanStdMeanStd
a22306371123290134294911205
b447813471994615206466616257
c24757111252317148325311215
d7159154928108381541010
e20565961158301237471513711024
f8677160835551204311010
g548013533322809654891010

Table 3 records means and standard deviations of the number of visited models in the 50,000 recorded drawingsofthechaininmodel space. Giventhatthemodel thatgenerated thedataisoneofthe , 768 possible models examined, we would want this to be as small as possible. For it is clear that the sample information is rather weak, allowing the chain to wander around and visit many models: as much as around a quarter of the total amount of models for prior f, and never less than six percent on average (Prior a). The sampler visits less models as n increases, and for n = 1000 we already have very few visited models for priors f in particular and also for d. When 10,000 observations are available, that is enough to make the sampler stick to one model (the correct one) for priors d, fand g. Surprisingly, whereas prior g still leads to very erratic behaviour ofthe sampler with , it never fails to put all the mass on the correct model for the larger sample sizes. Finally, note that even with 100,000observations, prior e still makes the sampler visit almost 110 models on average.

Table 4. Model 1: Means and Stds of Posterior Probabilities of Including each regressor.

Prior Reg. $\mathcal{N} = 50$
abcdefg
MeanStdMeanStdMeanStdMeanStdMeanStdMeanStdMeanStd
→10.980.070.980.070.980.060.970.090.980.070.960.090.980.06
20.220.170.290.140.240.160.330.100.210.170.350.090.360.13
30.250.180.310.140.270.170.350.110.240.180.370.100.380.13
40.270.190.330.150.280.180.370.110.250.190.390.100.400.13
→50.420.270.440.220.430.260.430.150.400.270.430.130.500.19
60.220.160.280.130.240.160.320.090.210.150.340.080.350.13
→70.940.140.940.130.940.140.900.150.940.150.870.150.940.12
80.220.160.290.140.240.150.320.100.210.150.340.080.360.13
90.210.140.280.120.230.150.320.080.200.140.340.070.350.11
100.210.140.280.120.230.140.320.080.200.130.340.070.350.11
→110.820.250.810.220.820.240.760.190.810.250.740.180.820.20
120.240.190.300.150.260.190.340.110.230.190.360.090.370.14
→130.390.270.430.230.400.260.440.180.380.270.450.150.490.20
140.270.220.320.190.280.210.360.130.250.220.370.110.390.16
150.220.150.280.120.230.150.330.080.210.150.350.070.360.11
$\mathcal{N} = 1000$
Prior Reg.abcdefg
MeanStdMeanStdMeanStdMeanStdMeanStdMeanStdMeanStd
→11.000.001.000.001.000.001.000.001.000.001.000.001.000.00
20.070.110.090.120.080.120.000.010.110.130.000.000.160.10
30.060.080.080.080.070.080.000.010.090.090.000.000.150.07
40.050.060.070.070.060.070.000.010.080.080.000.000.140.06
→51.000.001.000.001.000.001.000.001.000.000.880.261.000.00
60.070.080.090.090.070.080.000.000.100.100.000.000.160.08
→71.000.001.000.001.000.001.000.001.000.001.000.001.000.00
80.060.070.080.080.070.080.000.000.100.100.000.000.150.08
90.060.060.080.080.070.070.000.000.100.090.000.000.150.07
100.060.070.080.080.070.070.000.000.100.090.000.000.150.07
→111.000.001.000.001.000.001.000.001.000.001.000.001.000.00
120.060.100.080.100.060.100.000.010.090.100.000.000.150.09
→131.000.001.000.001.000.001.000.001.000.000.780.341.000.00
140.050.040.070.050.060.040.000.000.080.060.000.000.140.05
150.060.070.070.080.060.070.000.000.090.090.000.000.140.07

Table4indicatesinwhatsensethedifferentBayesianmodelstend toerriftheyassignposteriorprobability to alternative sampling models. In particular, Table 4 presents the means and standard deviations of the posterior probabilities of including each of the regressors. As we know from (4.9), Model 1 contains regressors 1,5,7,11 and 13 (indicated with arrows in Table 4). To save space, we shall only report this for and When , regressors and are almost always included. Since they are (almost) orthogonal to the other regressors, and their regression coefficients are rather large in absolute value, this is not surprising. Regressor is only correlated with and is still often included. The most difficult are regressors 5 and 13, which are positively correlated, and have relatively small regression coefficients ofopposite signs. The posterior probabilities ofincluding regressors notcontained in the correct model is relatively small. What is not clearly exemplified by Table 4 is that priors a through f tend to choose alternatives that are nested by Model 1 for small sample sizes, whereas prior g puts considerable posterior mass on models that nest the correct sampling model. Table 4 informs us that for the correct regressors are virtually always included. Only prior f has a tendency to choose models that are nested by Model 1. For the other priors there remain small probabilities ofincorrectly including extra regressors (the smallest for prior d and the largest for prior g). Alternative models tend to nest the correct model for all priors, except prior f, with this and larger sample sizes.

5.2.2. Results under Model 2

Let us now briefly present the results when the data are generated according to Model 2in (4.4), the null model. Table 5 presents means and standard deviations of the posterior probability of the null model. It is clear that this is not an easy task and most priors lead to small probabilities of selecting the correct model.

Overall, priors a and e do best for small sample sizes, whereas larger sample sizes are most favourable to priors a and b. Despite these rather small probabilities of the null model, the latter is still typically favoured over the second best model. This is evidenced by Table 6, where the three quartiles ofthe ratio of the posterior probabilities of Model 2 and the best other model are presented. Only prior f leads to a first quartile below unity. Overall, priors a and b seem to do best on this criterion. The difficulty of pinning down the correct (null) model can also be inferred from Table 7, where means and standard deviations of the number of visited models are presented. It is clear that some priors (like b, d, f and g) make the chain wander a lot for small sample sizes. Priors f and g retain this problematic behaviour even for sample sizes as large as 100,000. Interestingly, whereas prior fleads to (very slow) improvements as n increases, the bad behaviour with prior g seems entirely unaffected by sample size. Of course, we know from the theory in Subsection 3.2 that prior g does not lead to consistent Bayes factors in this case. The number of models visited is relatively small for prior a, which seems to emerge as the winner from the posterior results under Model 2.

Table 5. Model 2: Means and Stds of the posterior probability of the true model.

n50100100010,000100,000
PriorMeanStdMeanStdMeanStdMeanStdMeanStd
a0.03200.02690.07220.04940.38120.15430.71990.13460.89950.0529
b0.00210.00280.01140.01240.26060.15060.69100.14410.89600.0568
c0.00990.00810.02380.01570.14940.07150.41480.12010.70800.1009
d0.00030.00030.00060.00060.00660.00660.04270.03000.15700.0938
e0.04070.03220.07640.04920.22160.10500.35690.12200.47330.1235
f0.00010.00020.00020.00020.00050.00060.00090.00080.00140.0015
g0.00100.00120.00110.00120.00140.00140.00130.00110.00140.0014

Table 6. Model 2: Quartiles of ratio of posterior probabilities; True Model vs Best among the rest.

n50100100010,000100,000
PriorQ1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3
a2.44.47.03.16.29.26.016.124.830.060.086.278.7202.4272.1
b2.15.27.03.76.79.77.422.128.918.445.779.664.1153.8253.2
c1.31.82.11.22.22.72.04.17.16.315.221.816.140.370.2
d1.12.12.81.22.53.22.23.95.23.16.69.46.111.216.7
e2.45.37.53.56.89.34.911.417.58.115.725.312.926.936.2
f0.51.52.60.61.62.60.92.33.21.32.63.51.83.44.2
g1.22.33.21.32.53.21.32.73.21.72.73.51.62.53.1

Table 7. Model 2: Means and Stds ofNumber ofModels Visited.

n50100100010,000100,000
PriorMeanStdMeanStdMeanStdMeanStdMeanStd
a39219372213737445149116273710
b1448223791165115891813817245905417
c41837812496523513163159286012
d168692221157842097102512022485215761960690
e3529729239653994827951312732979
f179942105176012133163811879149633065133643671
g125871561125801483126531353125322201124132205

In summary, the posterior results for Model 1 point towards prior d as the best choice for most practical purposes, whereas prior a seems preferable for small samples and when the null model has generated the data.

5.3. Predictive inference

5.3.1. Results under Model 1

As discussed in Subsection 4.3 we shall condition our predictions on values of the regressors z . In all, we choose q = 19 different vectors for these regressors, and we shall focus especially on those vectors that lead to the minimum, median and maximum value for the mean of the sampling model. We shall denote these regressors as and , respectively. In our particular case, will be more extreme than zmax.

Table 8. Model 1: Conditional Medians of

$\eta$ 50100100010,000100,000
$z_{min}$ $z_{med}$ $z_{max}$ $z_{min}$ $z_{med}$ $z_{max}$ $z_{min}$ $z_{med}$ $z_{max}$ $z_{min}$ $z_{med}$ $z_{max}$ $z_{min}$ $z_{med}$ $z_{max}$
a2.4712.4252.4282.3912.3892.3912.3342.3552.3482.3262.3252.3382.3312.3502.345
b2.4802.4222.4312.4092.3902.3852.3342.3552.3472.3262.3252.3382.3312.3502.345
c2.4712.4242.4332.3972.3892.3892.3342.3552.3472.3262.3252.3382.3312.3502.345
d2.6912.4482.4752.5072.4062.4102.3582.3562.3542.3332.3262.3382.3322.3512.345
e2.4742.4282.4282.3932.3892.3912.3332.3552.3472.3262.3252.3382.3312.3502.345
f2.8362.4702.5302.6362.4232.4632.4752.3782.4122.4402.3552.3852.4172.3622.379
g2.4922.4302.4402.4502.3922.3952.4182.3622.3812.4192.3492.3742.4222.3642.382

Firstly, Table 8 presents the median of as computed in (4.7) across the 100samples conditionally upon the three vectors of regressors mentioned above. In interpreting these numbers, it is useful to recall that the theoretical minimum of the integral corresponding to is 2.335, as explained in Subsection 4.3. Of course, LPS in (4.7) is only a Monte Carlo approximation to this integral (based on a mere 100 drawings), so this lower bound is not always strictly adhered to. Under priors a through e we are predicting the sampling density virtually exactly with samples of size or more. Of these five priors, prior d performs slightly worse for very small samples (in particular, for the corresponding to the minimum sampling mean). Priors f and tend to be further from the actual sampling density and do not lead to perfect prediction even with 100,000observations. Prior f, in particular, leads to rather large values for LPS when n is 1000 or smaller and conditionally upon

Fig. 1. Model 1: Predictive densities,
Fig. 1. Model 1: Predictive densities,

In order to find out more about the differences between the predictive density in (4.5) and the sampling density in (4.3), we can overplot both densities for the three values of and . Figures 1 and 2 display this comparison for different values of n and the predictives for 25 of the 100 generated samples (to avoid cluttering the graphs). The dark line corresponds to the actual sampling density. Since the predictives from priors a-c and e are very close for all sample sizes, we shall only present the graphs for priors a, d, f and It is clear that for substantial uncertainty remains about the predictive distribution: different samples can lead to rather different predictives. They are, however fairly well calibrated in that they tend to lie on both sides of the actual sampling density for priors a, b, c and e and there is no clear tendency towards a different degree ofconcentration. These are exactly the priors for which takes on fairly small values (in between 0.02and 0.08). The priors d, fand g lead to much larger values for (in the range 0.16 to 0.41) and show a clear tendency for the predictive densities to be somewhat biased towards the median when conditioning on and . In addition, these priors induce predictives that are, on average, less concentrated than the sampling density. This behaviour can easily be understood once we realize that the locations of each of the components in (4.5) (the posterior mean of under each of the models considered) are clearly shrunk more towards the sample mean y¯ as becomes larger. This is, of course, in accordance with the zero prior mean for and the g-prior structure in (2.15). In addition, predictive precision decreases with , which explains the systematic excess spread ofthe predictives with respect to the sampling density for priors d, f and g. As sample size increases, the predictive distributions get closer and closer to the actual sampling distribution, and for or larger the effect of shrinkage due to has become negligible for prior d is then equal to 0.06) whereas it persists for priors f and g even with 100,000observations (where takes the values 0.14 and 0.16, respectively).

Fig. 2. Model 1: Predictive densities, n = 100, 1000, 10,000, and 100,000.
Fig. 2. Model 1: Predictive densities, n = 100, 1000, 10,000, and 100,000.

Table 9. Model 1: Medians of

n50100100010,000100,000
a2.4272.3822.3392.3342.333
b2.4272.3832.3392.3342.333
c2.4242.3812.3392.3342.333
d2.4732.4162.3472.3352.334
e2.4282.3822.3392.3342.333
f2.5022.4522.3932.3752.363
g2.4332.4522.3692.3662.366

We can also compare overall predictive performance, through considering for the 19 different values of and the 100 samples of . This leads to the results presented in Table 9, where the medians (computed across the 1900sample- combinations) are recorded for the different priors and sample sizes. Clearly, whereas all priors except for prior f (and, to a lesser extent, prior d) lead to comparable predictive behaviour for very small n, the fact that is constant in n makes prior lose ground with respect to the other priors as n increases. Prior falways performs worse than priors a through e. Note that priors a through e lead to median LPS values that are roughly equal to the theoretical minimum of 2.335, implying perfectly accurate prediction, for

Alternatively, we can compare the percentiles of the sampling distribution and the predictive in (4.5). We compute the predictive percentiles corresponding to the , and sampling percentile. The quartiles of these numbers, calculated over all 1900 sample-zf combinations, are presented in Table 10. This confirms that priors a-c and e lead to better predictions for small sample sizes, where the predictives from d,f and are too spread out. Starting at , prior d predicts well, whereas the inaccurate predictions with priors fand g persist even for very large sample sizes. Comparing Figure 2with Table 10, it is clear that most of the spread in the percentiles for priors f and with , 000 is due to the bias toward the median (shrinkage). Remember that Table 10averages over the 19different values of

5.3.2. Results under Model 2

As mentioned in Subsection 5.2.2, it is very hard to correctly identify the null model when we generate the data from such a model. On the other hand, prediction seems much easier than model choice. This can immediately be deduced from Table 10, where predictive percentiles are compared with the actual sampling percentiles. Clearly, the incorrectly chosen models are such that they do not lead our predictions (averaged over all the chosen models as in (4.5)) far astray. Table 10 contains predictive percentiles for and 1000, and even with just 50 observations median predictions are virtually exact, and the spread around these values is relatively small. Moreover, this behaviour is encountered for all priors. When sample size is up to 1000, prediction is near perfect for all priors.

Summarizing the predictive performance, we can state the following. For Model 1 priors a, b, c, and e (which induce the smallest values for seem to do best, but for 1000or more observations, prior d does just as well. All these priors predict virtually exactly for or more. Priors f and g imply too much shrinkage, as aresultoflarge values for (in the case ofprior g, this value does notdepend on n atall), and thus do worse in prediction. In the case of Model 2 (the null model), prediction is virtually perfect under all priors, even with small samples. For this model the issue of shrinkage is, of course, less problematic.

6. AN EMPIRICAL EXAMPLE: CRIME DATA

The literature on the economics of crime has been critically influenced by the seminal work of Becker (1968) and the empirical analysis of Ehrlich (1973, 1975). The underlying idea is that criminal activities are the outcome of some rational economic decision process, and, as a result, the probability of punishment should act as a deterrent. Raftery et al. (1997) have used the Ehrlich data set corrected by Vandaele (1978). These are aggregate data for 47U.S. states in 1960, which will be used here as well.

The single-equation cross-section model used here is not meant to be a serious attempt at an empirical study of these phenomena. For example, the model does not address the important issues of simultaneity and unobserved heterogeneity, as stressed in Cornwell and Trumbull (1991), but we shall use it mainly for comparison with the results in Raftery et al. (1997), who also treat it as merely an illustrative example.

We shall, thus, consider a linear regression model as in (1.1), where the dependent variable, y, groups observations on the crime rate, and the 15 regressors in Z are given by: percentage of males aged 14-24, dummy for southern state, mean years ofschooling, police expenditure in 1960, police expenditure in 1959, labour force participation rate, number of males per 1000 females, state population, number of nonwhites per 1000 people, unemployment rate of urban males aged 14-24, unemployment rate of urban males aged 35-39, wealth, income inequality, probability of imprisonment, and average time served in state prisons. All variables except for the southern dummy are transformed to logarithms.

In line with the recommendations from Section 5, we shall use prior a for this very small sample We run the chain to produce 100,000 draws after a burn-in of 25,000. This is more than enough to achieveconvergence, asisevidenced by thenearperfectcorrelation(0.9896) betweentheactual Bayesfactors computed as in (2.16) and the relative frequencies of model visits. All results will be based on the actual Bayes factors of the models visited (BARS, as explained in Subsection 5.1). In all, 3378 different models were visited, and the best 10% of those models account for 73.5% of the posterior model probability. Thus, posterior mass is not highly concentrated onjust a few models. Note that this run takes a mere 80seconds on a 120MHz 604 PowerPC Macintosh personal computer.

Table 11 presents the 9 models that receive over 1% posterior probability. The best model is the same as that in Raftery et al. (1997). In general, model probabilities are very similar, even though our prior is quite

Table 10. Quartiles of the Predictive Percentiles.

Model 1: Prior aModel 2: Prior a
$N = 50$ $N = 100$ $N = 1000$ $N = 10,000$ $N = 100,000$ $N = 50$ $N = 1000$
%Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3
.01.005.012.026.005.010.018.008.010.012.009.010.011.010.010.010.005.010.019.009.010.011
.05.024.052.094.029.049.074.045.050.058.048.050.052.049.050.050.031.048.073.046.050.054
.25.143.237.346.174.237.305.234.251.271.245.250.255.248.249.251.194.241.299.239.249.259
.50.339.471.600.391.479.565.480.502.526.494.500.507.498.500.502.432.496.559.489.498.510
.75.587.715.816.650.729.798.734.752.771.745.751.755.748.750.752.691.751.804.741.750.758
.95.871.932.965.906.941.963.944.951.957.948.950.952.950.950.951.922.949.969.947.950.954
.99.959.982.992.975.986.993.988.990.992.990.990.991.990.990.990.979.989.995.989.990.991
Model 1: Prior bModel 2: Prior b
$N = 50$ $N = 100$ $N = 1000$ $N = 10,000$ $N = 100,000$ $N = 50$ $N = 1000$
%Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3
.01.008.018.036.007.013.022.009.010.012.009.010.011.010.010.010.004.010.021.009.010.011
.05.032.064.107.035.054.079.045.051.058.048.050.052.049.050.050.027.047.079.045.050.054
.25.154.242.339.179.240.306.233.251.271.245.250.255.248.249.251.175.236.320.237.248.260
.50.333.453.569.385.469.549.479.501.524.494.500.507.498.500.502.405.491.586.487.498.512
.75.562.679.775.630.709.775.731.750.769.745.751.755.748.750.752.669.749.819.740.750.760
.95.836.903.944.891.927.953.943.950.956.948.950.952.950.950.951.916.950.972.946.950.954
.99.940.968.985.967.981.989.988.990.991.989.990.991.990.990.990.978.989.995.989.990.991
Model 1: Prior cModel 2: Prior c
$N = 50$ $N = 100$ $N = 1000$ $N = 10,000$ $N = 100,000$ $N = 50$ $N = 1000$
%Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3
.01.005.013.028.006.011.019.009.010.012.009.010.011.010.010.010.006.011.019.009.010.011
.05.025.053.096.030.050.075.045.051.058.048.050.052.049.050.050.032.049.074.046.050.054
.25.145.238.346.176.238.307.234.252.271.245.250.255.248.249.251.195.242.300.238.249.260
.50.339.470.601.391.479.565.479.502.525.494.500.507.498.500.502.432.496.559.488.498.510
.75.586.712.812.649.729.797.733.752.771.745.751.755.748.750.752.690.751.803.740.750.759
.95.868.929.963.905.939.962.944.951.957.948.950.952.950.950.951.922.949.968.946.950.954
.99.957.981.992.974.986.992.988.990.992.990.990.991.990.990.990.979.989.995.989.990.991
Model 1: Prior dModel 2: Prior d
$N = 50$ $N = 100$ $N = 1000$ $N = 10,000$ $N = 100,000$ $N = 50$ $N = 1000$
%Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3
.01.012.026.050.011.019.033.011.013.017.010.011.012.010.010.011.005.011.020.009.010.012
.05.040.075.126.041.064.099.049.058.067.050.052.055.050.051.051.029.049.077.044.050.055
.25.156.236.337.173.233.313.230.253.278.243.250.259.247.250.252.185.241.312.234.248.263
.50.310.420.539.351.436.530.458.486.517.485.495.505.495.498.501.419.494.573.480.498.516
.75.508.624.730.571.657.738.700.725.750.733.741.749.744.747.750.680.750.811.735.749.763
.95.777.855.911.836.885.926.924.934.944.942.945.948.948.949.950.920.949.970.945.950.955
.99.896.941.969.937.961.978.981.984.987.987.988.989.989.990.990.979.989.995.989.990.991
Model 1: Prior eModel 2: Prior e
$N = 50$ $N = 100$ $N = 1000$ $N = 10,000$ $N = 100,000$ $N = 50$ $N = 1000$
%Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3Q1Q2Q3
.01.005.012.026.006.010.018.009.010.012.010.010.011.010.010.010.006.011.019.009.010.011
.05.023.051.094.029.049.075.045.051.058.048.050.052.049.050.050.032.049.074.046.050.054
.25.143.236.346.174.237.306.233.251.271.245.250.256.248.249.251.196.241.299.238.248.260
.50.338.471.600.392.480.566.479.502.526.494.500.506.497.500.502.434.496.558.487.498.511
.75.587.715.817.650.730.798.732.751.770.744.750.755.748.750.752.691.751.803.740.750.759
.95.871.932.966.906.941.963.943.950.957.948.950.952.950.950.951.921.948.968.946.950.954
.99.960.982.992.975.986.993.988.990.992.989.990.991.990.990.990.979.989.995.989.990.991
Model 1: Prior fModel 2: Prior f
$N = 50$ $N = 100$ $N = 1000$ $N = 10,000$ $N = 100,000$ $N = 50$ $N = 1000$

Table 11. Models with more than 1% posterior probability.

Prob.Included Regressors
12.55%1349111314
22.48%1349111314
31.68%1359111314
41.52%134111314
51.41%13489111314
61.28%13491314
71.11%134911121314
81.04%1359111314
91.02%135111314

different from the one proposed in Raftery et al. (1997). In particular, we only require the user to choose the function and choosing it in accordance with our results in Section 5 leads to results that are very close to those with the rather carefully and laboriously elicited prior ofRaftery et al. (1997). In addition, the latter prior depends on the data, as mentioned in Subsection 2.3.

Table 12. Posterior Probabilities of Including each Regressor.

RegressorProb.
1Percentage of males age 14–2485.95 %
2Indicator variable for southern state22.46 %
3Mean years of schooling98.77 %
4Police expenditure in 196066.78 %
5Police expenditure in 195941.58 %
6Labor force participation rate14.74 %
7Number of males per 1,000 females15.24 %
8State population32.67 %
9Number of nonwhites per 1,000 people68.59 %
10Unemployment rate for urban males, age 14–2420.29 %
11Unemployment rate for urban males, age 25–3960.63 %
12Wealth30.79 %
13Income inequality99.94 %
14Probability of imprisonment90.73 %
15Average time served in prisons33.10 %
Figura

Fig. 3. Posterior density functions: Regressors 14 and 15.

Fig. 3. Posterior density functions: Regressors 14 and 15.

PosteriorprobabilitiesofincludingeachoftheregressorsaregiveninTable12, whichclearly indicatesthat schooling and inequality are virtually always included, while the percentage of males aged 14–24 and the probability ofimprisonment are also typically part ofthe relevant models. Overall, Table 12roughly agrees with Table 4 in Raftery et al. (1997). The deterrence variables are probability of imprisonment and average time served in prisons. These variables are ofparticular interest for the economic theory ofcrime, and their posterior density functions (averaging over models with posterior probabilities) are given in Figure 3. The coefficients of these regressors can be interpreted as elasticities. The gauge on top indicates (in black) the posterior probability of inclusion. The probability of imprisonment seems to have a moderately negative influence, as expected. The average time served in prisons, however, only has a posterior probability of

Fig. 4. Q-Q plot with 75%–25% sample split.

Fig. 4. Q-Q plot with 75%–25% sample split.

inclusion of one third (see also Table 13).

Ifwe split the data randomly into 35 observations used for estimation and 12to be predicted, we obtain the predictive Q-Q plot in Figure 4. This is what Raftery et al. (1997) call a “calibration plot”. We note that the best single model (graph in grey) predicts considerably worse than the predictive in (4.5) resulting from BMA (graph in black). Averaging over models with posterior probabilities does a much betterjob at predicting the 12 remaining observations than simply taking the model with highest posterior probability.

7. RECOMMENDATIONS

The prior structure we have proposed in Section 2only requires the choice ofone scalar hyperparameter, called . We make a possible function ofthe sample size and ofthe number ofregressors in the model under consideration, . Theoretical results on consistency (in the sense ofcorrectly identifying the model that generated the data ifthat model is contained in model space) suggest making a decreasing function ofsample size n. In addition, empirical results on posterior model choice and predictive performance seem to indicate that the following two priors are reasonable choices:

• prior a, where , for small n and data generated from models with relatively few regressors

• prior d, where , in other cases.

Thus, we would recommend the prior structure introduced here, together with these choices of for the purposes ofmodel selection or model averaging in linear regression models, whenever substantial prior information is lacking or a default analysis is the aim.

Our empirical simulation results compare favourably to those reported in George and McCulloch (1993) and Raftery et al. (1997), whereas our prior does not depend on the response variable and is very easy to elicit.

APPENDIX: SOME ASYMPTOTIC RESULTS

Before examining consistency, we need to establish some preliminary results. Some of these results are not new, whereas others are easy to derive. Thus, we do not present the proof of Lemma 1.

Lemma 1. Under the sampling model in (3.1),

(i) is nested within or is equal to model

\[\underset {n \to \infty} {\mathrm{plim}} \frac {y ^ {\prime} M _ {X _ {j}} y}{n} = \sigma^ {2}.\tag{A.1}\]

(ii) Under the assumption that for any model that does not nest ,

\[\lim _ {n \rightarrow \infty} \frac {\left(\alpha , \beta_ {s} ^ {\prime}\right) X _ {s} ^ {\prime} M _ {X _ {j}} X _ {s} \left(\alpha , \beta_ {s} ^ {\prime}\right) ^ {\prime}}{n} = b _ {j} \in (\mathbf {0}, \infty),\tag{A.2}\]

we obtain

\[\underset {n \to \infty} {\mathrm{plim}} \frac {y ^ {\prime} M _ {X _ {j}} y}{n} = \sigma^ {2} + b _ {j}.\tag{A.3}\]

A.1. Proof of Theorem 1

Denoting by the product of the first two factors in (2.16), we have that

\[C _ {j s} = \left(\frac {w _ {1} (k _ {j})}{g _ {0 j} + 1}\right) ^ {k _ {j} / 2} \left(\frac {g _ {0 s} + 1}{w _ {1} (k _ {s})}\right) ^ {k _ {s} / 2} w _ {2} (n) ^ {(k _ {s} - k _ {j}) / 2},\tag{A.4}\]

and thus

\[\lim _ {n \to \infty} C _ {j s} = \left\{ \begin{array}{l l} 0 & \text { if } k _ {j} > k _ {s} \\ 1 & \text { if } k _ {j} = k _ {s} \\ \infty & \text { if } k _ {j} < k _ {s}. \end{array} \right.\tag{A.5}\]

On the other hand, the limiting behaviour of the last factor in (2.16), which we denote by , depends on whether is nested within . We therefore consider the following three situations:

A.1.1. is not nested within and

Applying (A.3) we obtain

\[\underset {n \to \infty} {\operatorname{plim}} D _ {j s} = \underset {n \to \infty} {\lim} \left(\frac {\sigma^ {2}}{\sigma^ {2} + b _ {j}}\right) ^ {(n - 1) / 2} = \mathbf {0},\tag{A.6}\]

which, in combination with (A.5), leads directly to a zero limit for

A.1.2. is not nested within and

In this case, combining (A.5) with (A.6) no longer leads directly to the limit of . A natural sufficient condition leading to a zero limit for is given in (3.4), which ensures that converges to unity.

A.1.3 is nested within

Since in this case , we know from (A.5) that converges to zero. However, the limit of is now difficult to assess. Here we shall present sufficient conditions for a zero limit of . Rewriting as

\[D _ {j s} = \left(\frac {y ^ {\prime} M _ {X _ {s}} y}{y ^ {\prime} M _ {X _ {j}} y}\right) ^ {(n - 1) / 2} \left(1 + \frac {w _ {2} (n) \{w _ {1} (k _ {s}) (A _ {s} - 1) - w _ {1} (k _ {j}) (A _ {j} - 1) \} + w _ {1} (k _ {s}) w _ {1} (k _ {j}) (A _ {s} - A _ {j})}{\{w _ {2} (n) + w _ {1} (k _ {s}) \} \{w _ {2} (n) + w _ {1} (k _ {j}) A _ {j} \}}\right) ^ {(n - 1) / 2},\tag{A.7}\]

where

\[A _ {s} = \frac {(y - \overline {{y}} \iota_ {n}) ^ {\prime} (y - \overline {{y}} \iota_ {n})}{y ^ {\prime} M _ {X _ {s}} y} \mathrm{and} A _ {j} = \frac {(y - \overline {{y}} \iota_ {n}) ^ {\prime} (y - \overline {{y}} \iota_ {n})}{y ^ {\prime} M _ {X _ {j}} y},\]

itisimmediatethatthefirstfactorin(A.7) convergesindistributiontoexp , whereS hasa distribution with s degrees offreedom. On the other hand, the condition in (3.5) ensures afinite limitfor the second factor. Alternatively, if (3.6) holds, the second factor is (A.7) is smaller than one. Thus, using the fact that converges to zero, (3.5) and (3.6) each provide a sufficient condition for a zero limit of

A.2. Proof of results in Subsection

A.2.1. when is nested within

Since is nested within we have that . As a consequence, having a zero limit for the Bayes factor requires that w(·) be an increasing function, since otherwise . Provided that verifies this property, we obtain

\[\underset {n \to \infty} {\text { plim }} D _ {j s} = \underset {n \to \infty} {\text { lim }} \left(\frac {\sigma^ {2} + \frac {g _ {0 s}}{g _ {0 s} + 1} b}{\sigma^ {2} + \frac {g _ {0 j}}{g _ {0 j} + 1} b}\right) ^ {(n - 1) / 2} = \mathbf {0},\tag{A.8}\]

where b denotes the value in (A.2) corresponding to . This immediately leads to (3.7).

A.2.2. plim when is not nested within

Applying (A.1) − (A.3) and assuming that

\[\frac {g _ {0 s}}{g _ {0 s} + 1} b < \frac {g _ {0 j}}{g _ {0 j} + 1} b + \frac {1}{g _ {0 j} + 1} b _ {j},\tag{A.9}\]

where b corresponds to the model withjust the intercept and to in (A.2), we obtain that

\[\underset {n \to \infty} {\text { plim }} D _ {j s} = \underset {n \to \infty} {\text { lim }} \left(\frac {\sigma^ {2} + \frac {g _ {0 s}}{g _ {0 s} + 1} b}{\sigma^ {2} + \frac {g _ {0 j}}{g _ {0 j} + 1} b + \frac {1}{g _ {0 j} + 1} b _ {j}}\right) ^ {(n - 1) / 2} = \mathbf {0}.\tag{A.10}\]

Therefore, (A.9) ensures a zero limit for , and we can deduce the results presented in Subsection 3.2 for this case.

References

  1. Akaike, H. (1981), “Likelihood of a Model and Information Criteria,” Journal of Econometrics, 16, 3-14.

References

  1. Amemiya, T. (1986), Advanced Econometrics, Blackwell, Oxford.

References

  1. Atkinson, A.C. (1981), “Likelihood Ratios, Posterior Odds and Information Criteria,” Journal ofEconometrics, 16, 15-20.

References

  1. Bauwens, L. (1991), “The “Pathology” of the Natural Conjugate Prior Density in the Regression Model,” Annales d’Economie et de Statistique, 23, 49-64.

References

  1. Becker, G.S. (1968), “Crime and Punishment: An Economic Approach,” Journal of Political Economy, 76, 169-217.

References

  1. Berger, J.O. and Pericchi L.R. (1996), “The IntrinsicBayes Factor for Model Selection and Prediction,” Journal oftheAmerican Statistical Association, 91, 109-122.

References

  1. Bernardo, J.M. (1980), “A Bayesian Analysis ofClassical Hypothesis Testing,” (with discussion) in Bayesian Statistics, eds. J.M. Bernardo, M.H. DeGroot, D.V. Lindley and A.F.M. Smith, Valencia: University Press, pp. 605-618.

References

  1. Box, G.E.P. (1980),“Samplingand Bayes’ InferenceinScientificModellingand Robustness,” (withdiscussion) Journal oftheRoyal Statistical Society, Ser. A, 143, 383-430.

References

  1. Chib, S. and Greenberg, E. (1995), “Understanding the Metropolis-Hastings Algorithm,” The American Statistician, 49, 327-335.

References

  1. Chow, G.C. (1981), “A Comparison of the Information and Posterior Probability Criteria for Model Selection,” Journal ofEconometrics, 16, 21-33.

References

  1. Cornwell, C. and Trumbull, W.N. (1994), “Estimating the Economic Model of Crime With Panel Data,” Review ofEconomics and Statistics, 76, 1994, 360-366.

References

  1. Dawid, A.P. (1984), “Statistical Theory: The Prequential Approach,” Journal of the Royal Statistical Society, Ser. A, 147, 278-292.

References

  1. Dawid, A.P. (1986), “Probability Forecasting,” in: Encyclopedia of Statistical Sciences, Vol. 7, eds. S. Kotz, N.L. Johnson, and C.B. Read, New York: Wiley, pp. 210-218.

References

  1. Draper, D. (1995), “Assessment and Propagation of Model Uncertainty,” (with discussion) Journal of the Royal Statistical Society, Ser. B, 57, 45-97.

References

  1. Ehrlich, I. (1973), “ParticipationinIllegitimateActivities: A Theoretical and Empirical Investigation,” Journal of Political Economy, 81, 521-567.

References

  1. Ehrlich, I. (1975), “The Deterrent Effect of Capital Punishment: A Question of Life and Death,” American EconomicReview, 65, 397-417.

References

  1. Freedman, D.A. (1983), “A Note on Screening Regressions,” The American Statistician, 37, 152-155.

References

  1. Geisser, S. and Eddy, W.F. (1979), “A Predictive Approach to Model Selection,” Journal of the American Statistical Association, 74, 153-160.

References

  1. George, E.I. (1997), “Bayesian Model Selection,” Encyclopedia of Statistical Sciences, New York: Wiley.

References

  1. George, E.I. and McCulloch, R.E. (1993), “Variable Selection via Gibbs Sampling,” Journal of the American Statistical Association, 88, 881-889.

References

  1. George, E.I. and McCulloch, R.E. (1997), “Approaches For Bayesian Variable Selection,” Statistica Sinica, 7, 339-373.

References

  1. Gelfand, A.E. and Dey, D.K. (1994), “Bayesian Model Choice: Asymptotics and Exact Calculations,” Journal of the Royal Statistical Society, Ser. B, 56, 501-514.

References

  1. Geweke, J. (1996), “Variable Selection and Model Comparison in Regression,” in Bayesian Statistics 5, eds. J.M. Bernardo, J.O. Berger, A.P. Dawid and A.F.M. Smith, Oxford: Oxford University Press, pp. 609-620.

References

  1. Good, I.J. (1952), “Rational Decisions,” Journal oftheRoyal Statistical Society, Ser. B, 14, 107-114.

References

  1. Green, P.J. (1995), “Reversible Jump Markov Chain Monte Carlo Computation and Bayesian Model Determination,” Biometrika, 82, 711-732.

References

  1. Hannan, E.J. and Quinn, B.G. (1979), “The Determination of the Order of an Autoregression,” Journal of the Royal Statistical Society, Ser. B, 41, 190-195.

References

  1. Kass, R.E. and Raftery, A.E. (1995), “Bayes Factors,” Journal of the American Statistical Association, 90, 773-795.

References

  1. Kass, R.E. and Wasserman, L. (1995), “A Reference Bayesian Test for Nested Hypotheses and its Relationship to the Schwarz Criterion,” Journal oftheAmerican Statistical Association, 90, 928-934.

References

  1. Laud, P.W. and Ibrahim, J.G. (1995), “Predictive Model Selection”, Journal of the Royal Statistical Society, Ser. B, 57, 247-262.

References

  1. Laud, P.W. and Ibrahim, J.G. (1996), “Predictive Specification ofPrior Model Probabilities in Variable Selection”, Biometrika, 83, 267-274.

References

  1. Leamer, E.E. (1978), Specification Searches: Ad Hoc InferencewithNonexperimental Data, New York: Wiley.

References

  1. Lee, H. (1996), “Model Selection for Consumer Loan Application Data,” mimeo, Carnegie Mellon University.

References

  1. Madigan, D., Gavrin, J. and Raftery, A.E. (1995), “Eliciting Prior Information to Enhance the Predictive Performance of Bayesian Graphical Models,” Communications in Statistics, Theory and Methods, 24, 2271- 2292.

References

  1. Madigan, D. and York, J. (1995), “Bayesian Graphical Models for Discrete Data,” International Statistical Review, 63, 215-232.

References

  1. Min, C. and Zellner, A. (1993), “Bayesian and Non-Bayesian Methods for Combining Models and Forecasts with Applications to Forecasting International Growth rates,” Journal ofEconometrics, 56, 89-118

References

  1. Osiewalski, J. and Steel, M.F.J. (1993), “Regression Models Under Competing Covariance Structures: A Bayesian Perspective,” Annales d’Economieet deStatistique, 32, 65-79.

References

  1. O’Hagan, A. (1995), “Fractional Bayes Factors for Model Comparison,” (with discussion) Journal oftheRoyal Statistical Society, Ser. B, 57, 99-138.

References

  1. Pericchi, L.R. (1984), “An Alternative to the Standard Bayesian Procedure for Discrimination Between Normal Linear Models,” Biometrika, 71, 575-586.

References

  1. Phillips, P.C.B., (1995), “Bayesian Model Selection and Prediction With Empirical Applications,” (with discussion) Journal ofEconometrics, 69, 289-365.

References

  1. Poirier, D. (1985), “Bayesian Hypothesis Testing in Linear Models With Continuously Induced Conjugate Priors Across Hypotheses,” in Bayesian Statistics 2, eds. J.M. Bernardo, M.H. DeGroot, D.V. Lindley, and A.F.M. Smith, New York: Elsevier, pp. 711-722.

References

  1. Poirier, D. (1988), “Frequentist and Subjectivist Perspectives on the Problem of Model Building in Economics,” (with discussion) Economic Perspectives, 2, 121-144.

References

  1. Poirier, D. (1996), “Prior Beliefs About Fit,” in Bayesian Statistics 5, eds. J.M. Bernardo, J.O. Berger, A.P. Dawid and A.F.M. Smith, Oxford: Oxford University Press, pp. 731-738.

References

  1. Raftery, A.E. (1996), “Approximate Bayes Factors and Accounting for Model Uncertainty in Generalised Linear Models,” Biometrika, 83, 251-266.

References

  1. Raftery, A.E., Madigan, D. and Hoeting, J.A. (1997), “Bayesian Model Averaging for Linear Regression Models,” Journal oftheAmerican Statistical Association, 92, 179-191.

References

  1. Richard, J.F. (1973), Posterior and Predictive Densities for Simultaneous Equation Models, New York: Springer.

References

  1. Richard, J.F. and Steel, M.F.J. (1988), “Bayesian Analysis of Systems of Seemingly Unrelated Regression Equations Under a Recursive Extended Natural Conjugate Prior Density,” Journal of Econometrics, 38, 7-37.

References

  1. Schwarz, G. (1978), “Estimating the Dimension of a Model,” TheAnnals ofStatistics, 6, 461-464.

References

  1. Smith, A.F.M. and Spiegelhalter, D.J. (1980), “Bayes Factors and Choice Criteria for Linear Models,” Journal oftheRoyal Statistical Society, Ser. B, 47, 213-220.

References

  1. Vandaele, W. (1978), “Participation in Illegitimate Activities; Ehrlich Revisited,” in Deterrence and Incapacitation, eds. A. Blumstein, J. Cohen and D. Nagin, Washington D.C.: National Academy ofSciences Press, pp. 270-335.

References

  1. Zellner, A. (1986), “On Assessing Prior Distributions and Bayesian Regression Analysis With g-Prior Distributions,” in Bayesian Inferenceand Decision Techniques: Essays in Honour ofBrunodeFinetti, eds. P.K. Goel and A. Zellner, Amsterdam: North-Holland, pp. 233-243.

References

  1. Zellner, A. and Siow, A. (1980), “Posterior Odds Ratios for Selected Regression Hypotheses,” (with discussion) in Bayesian Statistics, eds. J.M. Bernardo, M.H. DeGroot, D.V. Lindley and A.F.M. Smith, Valencia: University Press, pp. 585-603.