Fundación de Estudios de Economía Aplicada
A Proposal to Distinguish State Dependence and Unobserved Heterogeneity in Binary Brand Choice Models by Jose Mª Labeaga-Azcona Mercedes Martos-Partal DOCUMENTO DE TRABAJO 2007-02
February 2007
We would like to thank Nora Lado and James Nelson for their very helpful comments and ACNielsen Spain for providing the scanner panel data used in the analysis. This research was partially supported by Ministerio de Educación y Ciencia Dir. Gral de Investigación, Grants SEJ2004-00672 and SEJ2005-08793-CO4-04.
* FEDEA and UNED.
** Universidad de Salamanca.
Los Documentos de Trabajo se distribuyen gratuitamente a las Universidades e Instituciones de Investigación que lo solicitan. No obstante están disponibles en texto completo a través de Internet: http://www.fedea.es
These Working Paper are distributed free of charge to University Department and other Research Centres. They are also availabl through Internet: http://www.fedea.es
Jorge Juan, 46 28001 Madrid -España Tel.: +34 914 359 020 Fax: +34 915 779 575 infpub@fedea.es
Abstract
This paper uses binary choice models that specify four possible sources of observed regularity in the consumer brand choice decision over purchase occasion: namely, state dependence, observed and unobserved heterogeneity and correlation effects. The objective is to distinguish correctly among the effects of these four variables. The estimation method proposed is an alternative to the most commonly used estimation methods in marketing choice models. We consider that the alternative method appropriately controls for observed heterogeneity and unobserved heterogeneity correlated with the state dependence variable because of the way the state dependence variable is built. The model is used for the first time in marketing following the methodology proposed by Chamberlain (1984). A relationship for unobserved heterogeneity is specified, taking into account the correlation among unobserved heterogeneity and other choice determinants. In this way, we split the influence of household state dependence and tastes on brand choice. The findings are very conclusive. We find that because the individual effects and the covariates are correlated, traditional estimation methods cannot be used to split state dependence and unobserved heterogeneity. The proposed model is found to yield better measures of predictive performance than the conventional model. The results are found to be robust across categories of laundry detergent and have significant implications for marketing policy.
Key Words: binary brand choice models, state dependence, unobserved heterogeneity, correlated effects, laundry detergent.
1. Introduction
Despite many advances in marketing brand choice models, the Guadagni and Little (1983) model still serves as the benchmark. In this, the key variable that allows the model to fit the choice probability accurately is the state dependence variable (sometimes called purchase feedback or loyalty). Since this seminal study, a rich literature on disaggregate brand choice models has emerged, supporting the existence of state dependence in the household’s brand choice decisions.
The influence of observed past experience (through actual purchase) with a brand on current choice probabilities is often referred to as structural state dependence (Heckman 1981). Put differently, if an identical household with no previously experienced event has different future behavior from a household with a previously experienced event, previous experience is the determinant for temporal unobserved persistence in brand choice (Hsiao 1986).
From a managerial perspective, strong state dependence effects imply a managerial incentive for inducing promotion (e.g., product sampling or price promotion). This is because if state dependence is present in household brand choice behavior, some households who bought the brand through price promotion or product sampling will be persuaded to stay with the brand after the promotion ends.
State dependence in brand choice can be influenced by chance in the market; that is, future event probabilities are disturbed by variables like price or promotions that change future brand choices. For example, low-income households can keep buying the cheapest choice. Therefore, state dependence can be explained as a consequence of environmental effects.
Persistence of brand choice can also be due to unobserved effects, but because these are present in the individual information set, they have a strong impact on event probability in the future. Unobserved heterogeneity refers to (residual) interindividual variations in purchase behavior, which cannot be explained by the observed brand choice experiences. These variations can be intrinsic preferences that consumers have for brands, or different ways in which consumers respond to marketing stimuli. Unobserved effects are also referred to as household tastes or household heterogeneity in preferences and prices. In contemporary households’ scanner panel data, household tastes are not available.
In this case, if the model of consumer behavior of the household includes only unobserved preference heterogeneity, from a managerial perspective, there will be less incentive for inducing promotion than in the case of state dependence because when the promotion ends, the household will revert to the preferred brand.
State dependence, unobserved heterogeneity and environmental effects are three sources of regularity in households’ decisions over time. However, it can sometimes be difficult to split the relative importance of these effects on household decisions. When the data contain economic and sociodemographic variables, it is also possible to control for environmental effects. However, state dependence and unobserved heterogeneity are often undistinguished. In this way, the managerial conclusions inferred from wrongly identified parameters could be erroneous.
Following Heckman (1981), there is another source of possible persistence in brand choice. This is the influence of prior propensities to select a brand on current selection probabilities or habit persistence. These habitual purchase inclinations are manifested as serial correlations in the random component of the utility functions. If one does not control for this contingency, state dependence may have an inordinately large impact on future choices. This is because previous experience is functioning as a proxy for serial correlation (in the utility-maximizing alternatives) that influences current choices.
It is important to distinguish among the effects of unobserved heterogeneity, state dependence and correlation effects. If heterogeneity is present in the true model (because households included in household scanner panel data are heterogeneous) and we ignore heterogeneity by fitting a model with only permit state dependence, then the state dependence parameter will be overestimated (Heckman 1981). Accordingly, we will incorrectly conclude state dependence (spurious state dependence). If unobserved heterogeneity and state dependence are correlated, and we do not correctly control for correlation in the empirical analysis, then previous experience is the only determinant in future experience because previous experience is a proxy for temporal unobserved persistence. Less problematic is the opposite situation, where if state dependence is present in the true model, and we ignore state dependence fitting a model which only permits heterogeneity, then the value of heterogeneity in the population will be overestimated .
Few studies have generally controlled for the four notions of temporal dependence in models of dynamic brand choice behavior: state dependence, environmental effects and unobserved heterogeneity and correlation effects.
1 In both cases, we assume that the correlation between state dependence and heterogeneity is positive. While this is common in applied terms, theoretically the correlation can also be negative.
This is because previous works have shown that correlation effects matter little in frequently purchased categories after heterogeneity and state dependence have been properly accounted for (Roy et al. 1996; Keane 1997; Seetharaman 2004). In fact, only Erdem and Sun (2001) test whether individual effects are uncorrelated with the covariates. As covariates, they include the price and display of the bought brand. They find individual effects are uncorrelated with the covariates. The outcome of this body of work is that marketing researchers appear to have little interest in building brand choice models that control for serial correlation in the residuals.
On the other hand, it is common in the economic literature to find evidence of correlation between individual effects and covariates. For example, Jones and Labeaga (2003) model current household consumption of tobacco as a function of the current real price of tobacco, some relevant sociodemographic variables of the household and past tobacco consumption. They also allow for the possibility of unobserved individual effects that are correlated with past tobacco consumption (reflecting the individual’s propensity to be addicted). In this case, sociodemographic variables are included as covariates. However, it is uncommon in the marketing literature to include this kind of variable as an explanation for brand choice.
We suggest that in the way in which the state dependence variable is built, the state dependence effect and unobserved heterogeneity are correlated. Abramson et al. (2000, p. 424) argue that the underspecification of serial correlation has serious consequences for the parameters in that the presence of serial correlation caused the loyalty coefficient to inflate. These differing views among econometricians and marketing researchers have brought about the need to undertake research on the ways in which serial correlation in the residuals in brand choice models can be controlled to avoid biased parameters.
In this analysis, we use a richer specification model where economic and sociodemographic variables are included to help control for the environmental effects. We test whether individual effects are correlated with these covariates, and we find that the individual effects are indeed correlated. The novelty of the analysis is that we allow correlation by introducing some assumptions to distinguish between state dependence and unobserved heterogeneity effects.
However, the major difficulty to estimate isolatable state dependence and unobserved preference heterogeneity is due to the fact that dependent variable is a latent variable. In the random utility model, the reality is that we observe the event (a household chooses or does not choose a brand), but the choice decision is made after a brand utility comparison among alternatives. Because the decision is to choose a single brand, we only know the chosen brand (the brand with higher utility). This is very useful in estimating models that assume continuous distribution (logistic or normal) errors, but at the same time, we cannot transform above unobserved preferences. There is not currently model transformation that can remove unobserved components.
In this paper, we do not only attempt to test whether brand choice regularity is also found in Spain but also provide a step forward by proposing a way to control for state dependence, observed heterogeneity (environmental effects) and unobserved heterogeneity. Unobserved heterogeneity correlated with the state dependence variable and other variables is included in the model specification. We follow the methodology proposed by Chamberlain (1984) that can be applied to linear and nonlinear models; in particular, brand choice models. Using this methodology, if state dependence and individual effects are correlated, then a functional relationship can be established where the unobserved effects are split into unobserved and state dependence effects on brand choice. In this way, it is possible to specify a reduced form for the model. The reduced form parameters can be consistently estimated in each time period and, then it is possible to derive, in a second stage, the ‘parameters of interest’ to obtain the correct prediction of the unobserved latent variable. By controlling for unobserved heterogeneity, we obtain structural form parameters such that the state dependence variable becomes less significant.
The aim of this paper is very clear. Starting with a simple binary choice model with state dependence, and in order to account for the heterogeneity of the brand choice process, we propose estimation in some different contexts proposed in the literature. We test for state dependence and unobserved heterogeneity in the consumer brand choice model. We then introduce to the marketing literature a new parameter estimation technique for heterogeneous logit models. We assume that unobserved heterogeneity depends on observed characteristics in a simple way; thus, we try to distinguish between the influences that unobserved heterogeneity and state dependence have on the probability of choosing a brand. To obtain robust results, we use two different product categories. We intuitively test all the model results and formally establish findings about the influence on choice of state dependence, unobserved heterogeneity, environmental and correlation effects. We compare the estimation results, the fit and the predictive ability of the traditional maximum likelihood estimation procedure with the estimation technique put forward here and we find better predictive performance in the method proposed.
The remainder of the paper is organized in four sections. In Section 2, we provide a brief literature review. In Section 3, we specify the proposed model and other nested models specified in the literature. We discuss the data used to estimate the models and the results in Section 4. Finally, in Section 5, we present the findings and our conclusions and suggest some managerial implications.
2. Literature Review
The control of the previous choice experience can be accomplished in a number of different ways in brand choice models. The key variable that allows the Guadagni and Little (1983) model to fit data accurately is the loyalty variable, which is an exponential smoothing of past purchases. The loyalty variable confounds two effects: state dependence and household heterogeneity. Heterogeneity refers to the differences across households in brand preferences or market responses, and state dependence refers to the impact of past purchases on current preferences.
The Guadagni and Little measure of brand loyalty is able to track the differences in purchase behavior across consumers and over time, but it cannot properly distinguish among sources of variation in utility from heterogeneity (across households) and sources of variation because of nonstationarity (within households over time). Many studies show that the Guadagni and Little loyalty variable does not sufficiently capture consumer heterogeneity (e.g., Ortmeyer et al. 1991; Fader and Lattin 1993). As pointed out by Lattin (1987), by using a single loyalty term, one implicitly assumes that differences across consumers and differences over time contribute equally to the heterogeneity in the base level utility. If such an assumption is inappropriate, it could have a distorting effect on the choice model. To avoid this problem, other measures have been proposed that split the cross-sectional and longitudinal effects. Jones and Landwehr (1988) proposed a discrete choice model that split heterogeneity and state dependence (by measuring the latest purchasing behavior). Modified versions of the Guadagni and Little loyalty variable have also emerged (Krishnamurthi and Raj 1988; Ortmeyer et al. 1991; Erdem 1996 and Keane 1997).
To avoid the critique usually applied to the loyalty variable in Guadagni and Little (1983), we model both sources of the variation in utility separately; that is, we separate the variation associated with nonstationarity by using the latest purchasing behavior and account for heterogeneity by specifying household sociodemographic characteristics and controlling for any unobserved preferences and response heterogeneity.
Unobserved heterogeneity can be included in model specification in a number of different forms (fixed and random effects) and with different assumptions (random effects correlated with environmental effects, correlated with the other choice’s determinants or without correlation). Approaches to estimating the parameters for unobserved heterogeneity in models of consumer brand choice behavior can be grouped into two broad classes: (1) those that estimate parameters for each household; a fixed effects model (e.g., Jones and Landwehr 1988) and the hierarchical Bayesian approach (e.g., Rossi and Allenby 1993); and (2) those that assume that household parameters are distributed according to a probability distribution and estimate the parameters of that distribution (random effects models). One approach is the finite mixture or latent class model approach, which captures heterogeneity across households in the form of discrete support points (e.g., Kamakura and Russell 1989; Chintagunta et al. 1991; Bucklin and Gupta 1992). The second approach is the continuous mixture approach, which models heterogeneity in the form of continuous mixture distributions (e.g., Gönül and Srinivasan 1993; Erdem 1996; Keane 1997). There has been some debate in the literature as to whether the finite mixture or continuous approach is better. However, research by Andrews et al. (2002) suggests that both approaches are equally good at parameter recovery and predictive validity and that “… whether an analyst prefers to use models with continuous or discrete representations of consumer heterogeneity is a matter of opinion and personal preference”.
In order to account for heterogeneity in the brand choice process, Jones and Landwehr (1988) introduced an extension of a technique developed by Chamberlain (1984). Chamberlain proposed a conditional maximum likelihood estimation technique. In this technique, household-specific parameters are taken out of the likelihood expression by conditioning on their sufficient statistics. One disadvantage of using this technique is that it can only explain the choice behavior of households that change brands, not households that are absolutely brand loyal or never choose the brand. This estimation technique can only explain the variety brand choice behavior. In the conditional heterogeneous model, they assumed that households in the sample that always purchase the brand or never purchase the brand will continue this pattern of behavior for all purchases. Under this assumption, purchases do not add to the likelihood function and so are dropped from the data set in the conditional estimation procedure. Thus, the estimation sample contains only households that ‘switch brands. Fixed logit models have rarely been used. One alternative is to assume some probability distribution for the intercept and slope terms using the random effect or hierarchical Bayesian approaches. These approaches eliminate some of the undesirable assumptions of the fixed effects model.
State dependence, unobserved heterogeneity and environmental effects are three sources of regularity on the household’s decisions over time. The state dependence variable usually is based on the household’s previous brand choice;
therefore, unobserved heterogeneity is correlated with the state dependence. Researchers have tried to split the relative importance of these three effects on household decisions. They have often found evidence for true state dependence in the choice process, even after controlling for a rich heterogeneity structure (e.g., Keane 1997; Ailawadi et al. 1999; Varki and Chintagunta 2004). However, there are sufficient empirical tests to support the introduction of observed heterogeneity in the model limits the explanatory power of the state dependence variable in brand choice decision models. It is also clear that controlling for unobserved heterogeneity reduce the explanatory capability of the state dependence variable (Keane 1997). All of these generalizations in brand choice models seek to remove spurious state dependence.
In sum, recent work on state dependence and heterogeneity in the context of disaggregate panel data of consumer brand choices has found evidence of true state dependence in the choice process. On the other hand, we also find empirical evidence of a zero-order choice process in models that use aggregate data. For example, Bass and Wind (1995) argued that zero-order consumer brand-choice behavior is an empirical generalization, and Uncles et al. (1995) concluded that applications of the Dirichlet model have provided strong evidence for zero-order choice process.
To build empirical evidence, we test for state dependence in consumer brand choice. First, we apply a simple test suggested by Chamberlain (1978) to distinguish state dependence from heterogeneity and serial correlation. This approach was subsequently applied in marketing by Erdem and Sun (2001), who found strong evidence of state dependence. Chamberlain’s test consists of including lagged exogenous variables (but not lagged choices) in the utility specification while allowing for unobserved heterogeneity in taste and response parameter. This is because Chamberlain argued that the key distinction between heterogeneity and state dependence is the dynamic response to the exogenous variables. If true state dependence is present, lagged exogenous variables affect current choices because they affect lagged choices. However, if true state dependence is not present, lagged exogenous variables cannot affect current choices. Thus, a test for whether lagged exogenous variables are significant determinants of current choices is also a test for state dependence.
Nonetheless, there are three main drawbacks of Chamberlain’s test. First, the test cannot be used to make further distinctions with regard to state dependence, heterogeneity, and serial correlation. Second, Chamberlain’s test depends on the assumption of individual effects being uncorrelated with the covariates if a random-effects specification is used to implement the test (or the assumption that such correlation is correctly modeled). Third, Chamberlain’s test as a ‘test of state dependence’ requires at least one exogenous variable that would not have a lag structure in the absence of state dependence. Chamberlain (1984) proposed a new methodology to avoid these drawbacks.
We follow the methodology proposed by Chamberlain (1984) that has never before been applied to marketing. Using this methodology, if state dependence and unobserved heterogeneity are correlated, then a functional relationship can be established on the unobserved effects to separate them from the state dependence effects on brand choice. This methodology is not only a test of state dependence, but also a way to split state dependence and unobserved heterogeneity. Another advantage is that the variables included need not be strictly exogenous.
We use a random effects specification and estimate the structural parameters using the GMM procedure. Working in this way, we avoid a disadvantage of the conditional maximum likelihood technique in that it can only explain the choice behavior of households that change brands, not households that are absolutely loyal or households that never choose the brand. The technique used in the current analysis can explain any type of brand choice behavior, including variety and no-variety behavior.
3. Model Specification and Estimation Methods
3.1 General model
We specify a binary choice model that includes state dependence, environmental effects and unobserved heterogeneity in preferences and the price response. We express the model in the following equations:
\[y _ {i t} ^ {*} = \beta_ {i} p _ {i t} + \lambda y _ {i t - 1} + \delta z _ {i} + v _ {i t}\tag{1}\]
\[\nu_ {i t} = \alpha_ {i} + \mu_ {i t}\tag{2}\]
\[y _ {i t} ^ {*} = \alpha_ {i} + \beta_ {i} p _ {i t} + \lambda y _ {i t - 1} + \delta z _ {i} + \mu_ {i t}\tag{3}\]
\[y _ {i t} = \{1 \text { if } y _ {i t} ^ {*} > 0; 0 \text { whether } y _ {i t} ^ {*} \leq 0 \}\]
\[y _ {i t - 1} ^ {*} = \alpha_ {i} + \beta_ {i} p _ {i t - 1} + \lambda y _ {i t - 2} + \delta z _ {i} + \mu_ {i t - 1}\tag{4}\]
Where is the utility of the brand choice for consumer i at purchase occasion in the product category, is the price parameter for consumer i, and is the price of the chosen alternative for consumer i on purchase occasion t. In the literature, this is referred to as ‘unobserved price heterogeneity’ because it contains effects unobserved to the researcher. We assume that is a time-invariant effect. are socio-demographic variables specific to household i which do not exhibit time variation and is the vector of parameters associated with them. Information about allows knowledge of the households’ observed heterogeneity, λ is the state dependence parameter, , and is a lagged term of the dependent variable that incorporates purchase-event feedback. denotes an indicator about consumer i’s previous brand choice behavior on purchase occasion t–1, and is the householdspecific intercept term that characterizes the differences in brand preferences among consumers and is time invariant. In the literature, this is referred to as ‘unobserved preference heterogeneity’ because it also contains effects unobserved by the researcher. The specific household preference effect is obtained when we discompose the random error following equation (2) into a household preference effect and mixed error that changes both across time and households.
We further assume that the mixed error components are independent and identically Weibull distributed. Thus, the conditional probability of the chosen alternative at time t by consumer i is given by the binary logit model.
\[P \left(\mathrm{y} _ {\mathrm{it}} = 1\right) = \frac {\exp \left(\alpha_ {i} + \beta_ {i} p _ {i t} + \lambda \mathrm{y} _ {i t - 1} + \delta z _ {i}\right)}{1 + \exp \left(\alpha_ {i} + \beta_ {i} p _ {i t} + \lambda \mathrm{y} _ {i t - 1} + \delta z _ {i}\right)}\]
and
\[P \left(\mathrm{y} _ {\mathrm{it}} = 0\right) = \frac {1}{1 + \exp \left(\alpha_ {i} + \beta_ {i} p _ {i t} + \lambda y _ {i t - 1} + \delta z _ {i}\right)}\tag{5}\]
In general, we must assume some observability rule linking observed and latent components. The rule is contained in the second part of equation (3), and it establishes the relationship between (the utility of choosing the brand) and (the true choice realization), namely, the variable takes a value of one if the consumer i chooses the alternative on purchase occasion t. The difficulties in separating the state dependence effect and the unobserved preference effects have been proven sufficiently in the literature, and we reiterate the problem in this paper. However, the contribution of this paper is that we estimate the state dependence effect on brand choice when we control for the unobserved heterogeneity in the correct way. The proposed model is based on the specified general model; we also consider another nested model in the general model that we use to test the proposed model. We follow a sequential estimation process with the aim of differentiating each effect as a determinant on brand choice, and at the same time, we obtain an intuitive test for each group of variables.
3.2 Proposed model
In nonlinear models, like those applied to brand choice behavior, it is not possible to use transformation to first differences or orthogonal deviations to yield estimators of the λs that are asymptotically independent of the householdspecific effects and hence, consistent for all the parameters. A conditional likelihood approach can be followed in order to sweep out the fixed effects (Chamberlain 1984, pp. 1274–1278). However, the dynamic specification of equation (3) adds an additional difficulty in that the presence of lags of the latent endogenous variable induces correlation between this regressor and the effects. This is especially difficult to control for in nonlinear models.
There are some circumstances where λ could not be identified. First, could be correlated with such as shown in equation (4), in a way such as is at best predetermined for the mixed error but correlated with the time invariant part of the error. In this situation, λ cannot be separately identified from the correlation effect.
Second, it is possible that in different situations at the household level, the state dependence parameter cannot be separated from the household-specific effect: (i) in case of absolute loyalty, state dependence will be a constant for the household, as it is the household-specific effect, and therefore we cannot separate both effects; (ii) when the alternative is never chosen, it implies the absence of loyalty, then the state dependence for the household takes the value zero for all purchase occasions. Once again, we will have two constants and cannot separate the effects of both choice determinants.
One of the advantages of using panel data is the possibility of accounting for the correlation among the effects and the explanatory variables. Chamberlain’s (1984) suggestion of using a random effects approach and specifying a distribution for the effects conditional on the exogenous variables can be applied to (3).
We follow Chamberlain in assuming:
\[E \left(\alpha_ {i} / X _ {1}\right) = \sum_ {t = 0} ^ {T} \pi_ {t} ^ {\prime} X _ {1 i t} + \pi_ {r} ^ {\prime} R _ {i t} + w _ {i}\tag{6}\]
where are considered exogenous variables and contains nonlinear terms and interactions in We have to choose the instrument set carefully. Natural choices for instruments are past prices as well as demographic variables not included in the household decision set. However, we do not use past prices to avoid potential problems with multicollinearity. We expect these instruments to be highly correlated with the unobserved effects and uncorrelated with the mixed error2.
If we substitute (6) in (1), we obtain the reduced form of the general model in the following equation:
\[y _ {i t} ^ {*} = \pi W _ {i} + \varepsilon_ {i} \mathrm{i} = 1, \dots , \mathrm{N}\tag{7}\]
where we can derive the parameters of interest and λ) from the non-linear relationship between them and the reduced form parameters in (7) Chamberlain (1984) suggested that the parameters in the reduced form (7) could be estimated by fitting a model for each time period where we have information3. We suppose the unobserved effects to be time invariant, but this hypothesis is not very strong because the period used in the current analysis is short.
After having the estimator , we can build predictions for the latent variables, , and derive ‘parameters of interest’ and λ) in a second step, using the following equation.
\[\hat {y} _ {i t} ^ {*} = \alpha_ {i} + \beta_ {i} p _ {i t} + \lambda y _ {i t - 1} + \delta z _ {i} + u _ {i t}\tag{8}\]
Finally, we derive the relevant vector of parameters by applying GMM to (8).
In the discussion in this section, we have shown two assumptions concerning the state dependence effect on brand choice. The first assumption implies that it is the realization of the dependent variable in the previous period, , what influences the choice probability in the current period. In this way, state dependence is in the information set of the analyst in period t and as a result it is an observed variable. The alternative assumption implies that the choice brand probability in period t could be influenced by the previous choice probability; namely, the variable that influences the probability in t is . We can postulate it because of missing information to the analyst or because the current choice probability is also influenced by the probability of choosing the same alternative in t-1. If this were the case in (1), we would have instead of , and as consequence, in (8) we would also replace the latent explanatory variable by the predicted values using the reduced form predictions for
Xlit
2 We include in X1it the age of the main buyer of the household, age of the household’s wife, number of children of different ages and number of adults, dummies for sex, dummies indicating the geographic location of the household, and dummies indicating the population size of the geographic location of the household and household expenditure on the category during the analyzed period. We include in the age squared and age Rit cubed of the main buyer of the household, the age squared of the household’s wife, and interactions between the sociodemographic variables.
3 In the empirical application, we detail how the time period is defined.
In both cases, this approach assumes that the conditional expected value of the household preference effect is linear and independent between the variables in W and the disturbance . We assume that the mixed error in (8) does not present autocorrelation of any order. Using these assumptions, we derive the parameters of the structural form by any method that allows for the control of the household preference unobserved effects in equation (8). Chamberlain (1984) proposed deriving parameters using a minimum distance process; but we can also use alternative methods, as least squares or instrumental variables (on the first differences of the model).
Estimators of (8) are sensitive to assumptions about the distribution of the linearity of the expected value (6) and the conditional mean independence assumption implied by (8). However, these hypotheses can be checked by specification tests at the level of the reduced form. In order to estimate the model we make use of the fact that the distribution of conditional on the explanatory variables but marginal to the effects, is of the same form as the joint distribution (Chamberlain 1984, section 3.1). This allows us to estimate the reduced form (7) using discrete choice models for each t and then combining predictions as shown in (8).
3.3 Nested models
We compare the performance of the proposed model with the traditional brand choice model proposed by the literature. We compare our model with three models nested in the general one. These models are frequently used in the marketing literature.
Model 1
We consider that choice can be influenced by prices, state dependence and sociodemographic variables. We keep the assumption that unobserved household preference and price effects do not exist; therefore and This assumption implies that all the households have the same preferences toward the brand and the same response to prices. We can rewrite (5) as follows.
\[P \left(\mathrm{y} _ {\mathrm{it}} = 1\right) = \frac {\exp \left(\alpha + \beta p _ {i t} + \lambda \mathrm{y} _ {i t - 1} + \delta z _ {i}\right)}{1 + \exp \left(\alpha + \beta p _ {i t} + \lambda \mathrm{y} _ {i t - 1} + \delta z _ {i}\right)}\]
and
\[P \left(\mathrm{y} _ {\mathrm{it}} = 0\right) = \frac {1}{1 + \exp \left(\alpha + \beta p _ {i t} + \lambda y _ {i t - 1} + \delta z _ {i}\right)}\tag{5'}\]
This model is very similar in nature to the model in Guadagni and Little (1983). One critique is that under the presence of unobserved effects, the state dependence will be overstated if the heterogeneity is positively correlated with the choice determinants.
Model 2
We remove the assumption that unobserved preference and price effects are not present and consider that unobserved effects exist and they follow a random form specification. We keep equation (5).
\[P \left(\mathrm{y} _ {\mathrm{it}} = 1\right) = \frac {\exp \left(\alpha_ {i} + \beta_ {i} p _ {i t} + \lambda y _ {i t - 1} + \delta z _ {i}\right)}{1 + \exp \left(\alpha_ {i} + \beta_ {i} p _ {i t} + \lambda y _ {i t - 1} + \delta z _ {i}\right)}\]
and
\[P \left(\mathrm{y} _ {\mathrm{it}} = 0\right) = \frac {1}{1 + \exp \left(\alpha_ {i} + \beta_ {i} p _ {i t} + \lambda y _ {i t - 1} + \delta z _ {i}\right)}\tag{5}\]
The specification is estimated using the latent class model (Kamakura and Russell 1989; Wedel and Kamakura 1998). Model 1 and Model 2 can be estimated using traditional unconditional maximum likelihood estimation procedures.
Model 3
We specify Model 3 only to allow us to test for state dependence. We suppose price as the lagged covariate to be included in the model. If there is state dependence, a consumer is more likely to buy a brand on the current purchase occasion if they bought it on a previous occasion. Then, even though the lagged price has no effect on the current purchase decision, state dependence induces a negative correlation between the lagged price and the current purchase probability for a brand (Erdem and Sun 2001). Chamberlain’s (1978) approach entails testing for such correlation. We can rewrite in this case (5) as follows.
\[P \left(\mathrm{y} _ {\mathrm{it}} = 1\right) = \frac {\exp \left(\alpha_ {i} + \beta_ {i} p _ {i t - 1} + \delta z _ {i}\right)}{1 + \exp \left(\alpha_ {i} + \beta_ {i} p _ {i t - 1} + \delta z _ {i}\right)}\]
and
\[P \left(\mathrm{y} _ {\mathrm{it}} = 0\right) = \frac {1}{1 + \exp \left(\alpha_ {i} + \beta_ {i} p _ {i t - 1} + \delta z _ {i}\right)}\tag{5''}\]
This last specification is also estimated using the latent class model. We must take into account that Chamberlain’s test depends on the assumption of absence of correlation between the individual effects and the covariates. Therefore, if correlation is present, the test is incorrect.
One difficulty with these nested models is that we cannot transform (1) to remove the unobserved household effects. In fact, transformation is only possible in model (8) where we replace the latent variable using its predicted counterpart. The problem is that in maximum likelihood estimation, it is not possible to split the ‘parameters of interest’ from the unobserved household effects. Because we have few observations for each household, it is not possible to estimate the unobserved household effects consistently, and these inconsistencies are translated to the parameter of interest. This is called the incidental parameter problem (for more details, see Chamberlain 1984 or Bover and Arellano 1988). Conversely, these models do not assume correlated effects.
4. Estimation and Empirical Tests
In this section, we present the data and variables. We also present our results for the models specified above. In all cases, we use a binary discrete choice model, in particular the logit model, as a simple way to split state dependence and unobserved heterogeneity. We could also use models with normal errors, but the results are very similar. The third part of this section is used to compare the results.
4.1 Brand choice, data and variables
We define the dependent variable of the logit choice model by taking into account the store brand market share in Spain. Store brands, also known as private labels or retail brands, have enjoyed increased success in recent years. Europe shows a traditional dominance in terms of the market share of store brands. Spain is among the top five markets (ACNielsen 2005) where the market share of store brands reached 26% in 2005, and according to the ACNielsen forecast, sales will increase at double the rate of national brands in 2006. Laundry detergents and non-food groceries is one of the categories in which store brands have had more success in Spain, enjoying a market share approaching 30%. Therefore, the choice indicator takes a value of one when household i purchases a store brand (SB) on occasion t and zero otherwise (national brands, NBs).
We estimate our models using scanner panel data, supplied by ACNielsen Spain, on household purchases in two product categories: fine laundry detergent and non-fine laundry detergent. SBs are present in both categories. The Spanish data set includes a representative sample of households across the country, rather than households in specific cities. Their purchase activities are recorded from January 1999 to December 2000. We have 1,107 households accounting for 5,347 purchases of fine laundry detergent, and 1,557 households accounting for 33,246 purchases of non-fine detergent. To obtain robust results, we estimate the models in these two detergent categories because store brand market shares differ.
Households without information during the estimation period or during the prediction period were dropped from the analysis. We use the first-year period to estimate the models, while the second-year period is used to predict the choice market share. In the fine laundry detergent category, we have 622 households with 4,366 purchase occasions (2,172 in the estimation period and 2,194 in the prediction period) and 1,499 households with 30,050 purchase occasions for nonfine laundry detergent (15,402 in the estimation period and 14,648 in the prediction period).
To analyze the relative importance of state dependence compared with other variables in the model, we include some explanatory variables describing brands and consumers. The variables included in the model specifications are grouped into the following three categories.
Purchase-occasion-specific variable:
‘Priceit’ is the price/weight of the bought brand on purchase occasion t of household i. Weight is equal to one kilogram and price corresponds to the sales price4.
State dependence variable:
‘Lastit’ is a dummy variable that reflects the relative impact of recent choice behavior, measured by whether household i purchased the brand on occasion t–1 (Jones and Landwehr 1988, variable of purchase-event feedback).
Household-specific sociodemographic variables:
4 We only have shelf prices for the bought brands. Information about features and displays is not available; nor do we have data about the size discounts associated with price promotion.
‘Sizei’ is a count variable ranging from 1 to 5 representing the size of household i, where value 5 identifies households with five or more members.
‘Social class is a dummy variable that takes a value of one for high and high-medium class households and zero otherwise.
‘Social class is a dummy variable that takes a value of one for medium class households and zero otherwise.
‘Social class is a dummy variable that takes a value of one for mediumlow and low class households and zero otherwise.
‘Workeri’ is a dummy variable, where a value of one identifies households with a working housewife.
Table A.1 of the Appendix provides descriptive statistics for the dependent and explanatory variables for the two categories of detergent. When we compare the fine and the non-fine laundry detergent markets, we find that the former had a larger store brand market share, and the price gap between store brands and national brands is large. Households choose store brands on 37 percent of purchase occasions for fine laundry detergent and on 19 percent of purchase occasions for nonfine laundry detergent. The fine laundry detergent category is a less-frequently bought category.
Market concentration is higher in the fine laundry detergent category. In this market there are 61 bought brands: 23 SBs and 38 NBs. However, in the non-fine laundry detergent category, there are 116 bought brands: 39 SBs and 77 NBs5.
4.2 Estimation model, results and tests
We estimate and discuss the models specified in Section 3. The estimation results are shown in Table 1 for the case of fine laundry detergent and in Table 2 for the case of non-fine laundry detergent. All models are estimated in the firstyear period of purchase occasion by the household in both categories, and we keep the second-year period as the prediction sample.
To estimate the proposed model, we need to define a period where we can get a consistent estimator. We defined as the time unit the month of purchase for the non-fine laundry detergent and the quarter of purchase for the fine laundry detergent because fine laundry detergent is bought less frequently. We could choose other time units, but our choice is based on the interpurchase time in each category6.
According to Dhar and Hoch (1997) and Ailadawadi and Keller (2004), SBs appear to enjoy a higher share in large, less-promoted categories with high market concentration when the price gaps between national brands and store brands are large. We found evidence of this in the differing success of SBs in the Spanish laundry detergent market.
As remarked in Sections 1 and 3, the major problem in estimating isolatable state dependence and unobserved preference heterogeneity is that the dependent variable is a latent variable. Therefore, the model does not accept any transformation that can remove the unobserved components, which are very useful when using scanner or panel data. For this reason, after we specify a relationship between the individual preference effects and observed exogenous variables, such as that described in equation (6); we fit a binary logit model in each time period and obtain the predictions. At this point, it is still possible at the second stage to derive the structural parameter using methods where transformations are possible. In the second stage, the estimation method is random errors GMM because these methods allow us to model the random household effects from the model specification. In this way, coefficient estimates, including the corresponding to the state dependence variable, will be bias free.
We assume that there is the same heterogeneity structure in price and in preferences as in Model 2. We also adjust the probability of membership to each particular segment following Kamakura and Russell (1989) and Gupta and Chintagunta (1994); using an assignment rule such as ‘membership in the segment with highest probability’, it is possible to assign households uniquely to segments with differential preference and price sensitivity.
We analyze the results shown in Tables 1 and 2. We make comparisons of the t-statistics from the variables in Model 1 and Model 2 in each category. Because the t-statistic for the last variable is large, we conclude that state dependence is an important choice determinant for these two models in both categories. The results for Model 1 show evidence that in the brand choice decision the household’s previous choice (state dependence) increases the probability of choosing the same brand. In this model, state dependence is the most important brand choice determinant. However, if unobserved effects exist and are positively correlated with choice determinants, the effect of state dependence is overestimated. We check whether there is any unobserved heterogeneity. We use a latent class model with heterogeneity in preferences and prices. We use the four information criteria suggested by Elrod and Keane (1995) to choose the optimal latent class model. They are the Akaike Information Criterion (AIC), the Hannan–Quinn (HQ) Criterion, the Bayesian
6 In the full time period of 727 days under study, the median interpurchase occasions of fine laundry detergent is 8 purchases and 26 purchases of non-fine laundry detergent; therefore, we can approximate a time unit of one month for non-fine laundry detergent and one quarter for fine laundry detergent.
Information Criterion (BIC), and the Consistent Akaike Information Criterion (CAIC). The information criteria are computed as AIC=–2LogL+2K, HQ=– 2LogL+2Kln(Ln(N), BIC=–2LogL+Kln(N) and CAIC=–2LogL+K(ln(N)+1), where LogL is the value of the log-likelihood function for each model, K is the number of parameters estimated, and N is the sample size. We prefer those models with higher values of the log-likelihood and smaller values of AIC, HQ, BIC and CAIC. The values of LogL, K, N, AIC, HQ, BIC and CAIC for the optimal model in each category are reported as measures of fit in Tables 1 and 2.
Table 1. Estimation results for the logit model in the fine laundry detergent category
| Model 1 | Model 2 | Model 3 | Proposed model | ||
| (1-step) | (2-step) | ||||
| VARIABLES | |||||
| Last | 2.71 (23.39)*** | 1.16 (5.93)*** | 0.58 (2.88)*** | ||
| Price | -0.73 (-4.12)*** | -77.11 (-6.01)***5.68 (5.73)*** -3.90 (-6.94)*** | -5.39 (-70.50)*** -2.51 (-19.38)*** -1.96 (-33.26)*** | ||
| Lagged price | -1.61 (-3.67)** -1.26 (-1.28)* 0.44 (0.79) | ||||
| Worker | -0.26 (-2.31)** | -0.38 (-1.72)** | -1.01 (-3.15)*** | -0.12 (-0.63) | |
| Size | -0.04 (-0.95) | 0.18 (1.82)** | 0.28 (2.25)** | 0.04 (0.43) | |
| Social Class 1 | 0.10 (0.64) | -0.61 (-1.83)** | 0.38 (0.92) | 1.18 (3.84)*** | |
| Social Class 2 | 0.06 (0.52) | -0.16 (-0.57) | 0.83 (2.53)** | 0.69 (2.70)*** | |
| Dummy of segments 1 | -1.08 (-3.15)*** | ||||
| Dummy of segments 2 | -0.41 (-1.32) | ||||
| Constant | -0.42 (-1.48)* | 56.01 (6.01)*** -7.51 (-5.94)*** 4.04 (4.87)*** | 0.76 (1.10) 3.12 (2.49)*** -5.50 (-5.66)*** | 6.67 (13.55)*** | |
| MEASURES OF FIT | |||||
| Log likelihood | -1,087 | -857 | -662 | -400 | |
| # of purchases | 2,172 | 2,172 | 1,550 | 2,172 | 2,172 |
| # of segments(s) | S = 1 | S1 = 0.46 S2 = 0.18 S3 = 0.36 | S1 = 0.32 S2 = 0.16 S3 = 0.52 | S1 = 0.46 S2 = 0.18 S3 = 0.36 | |
| # of parameters | 7 | 13 | 12 | 70 | 11 |
| AIC | 2,188 | 1,740 | 1,348 | 1,360 | |
| HQ | 2,203 | 1,767 | 1,372 | 1,942 | |
| BIC | 2,282 | 1,914 | 1,500 | 5,103 | |
| CAIC | 2,237 | 1,829 | 1,426 | 3,233 | |
| Pseudo R2 | 0.70; 0.71; 0.73; 0.70 | ||||
| Notes:1. Models:Model 1. Logit model with state dependence, observed heterogeneity and without unobserved heterogeneity.Model 2. Latent class logit model with state dependence and observed heterogeneity.Model 3. Latent class logit model without state dependence and with observed heterogeneity and lagged price.Proposed Model. Panel model with state dependence and observed heterogeneity controlling unobserved heterogeneity using random effect. Model with two stages.2. Table entries are the value coefficient with the t-statistic in parentheses, *p < 0.10, **p < 0.05 and ***p < 0.01.3. Akaike Information Criterion (AIC = -2LogL + 2K); Hannan–Quinn (HQ = -2LogL + 2Kln(Ln(N)); Bayesian Information Criterion (BIC = -2LogL + Kln(N)); and Consistent Akaike Information Criterion (CAIC = -2LogL + K(ln(N) + 1)). Here, LogL is the value of the log-likelihood function for each model, K is the number of parameters estimated, and N is the sample size.4. Pseudo-R2 = 1 - (LogL/ LogLC), where LogL is the value of the log-likelihood function at the optimum for the complete model and LogLC is the value of the log-likelihood function at the optimum for a restricted model with only a constant as the explanatory variable.5. Variables used in the first step are the age of the main buyer of the household, age of the wife in the household, number of children of different ages and number of adults, dummies for sex, dummies indicating the geographic location of the household and dummies indicating the population size of the geographic location of the household, and household expenditure on the category during the analyzed period. Age squared and age cubed of the main buyer of the household, age squared of the household wife, and interactions between the sociodemographic variables. | |||||
Table 2. Estimation results for the logit model in the non-fine laundry detergent category.
| Model 1 | Model 2 | Model 3 | Proposed model | ||
| (1-step) | (2-step) | ||||
| VARIABLES | |||||
| Last | 2.95 (59.91)*** | 1.50 (19.40)*** | 0.34 (6.68)*** | ||
| Price | -1.15 (-10.75)*** | -33.47 (-19.65)*** | -3.70 (-138.4)*** | ||
| 7.46 (13.99)*** | -1.89 (-62.50)*** | ||||
| -0.11 (-0.76) | -1.22 (-47.85)*** | ||||
| Lagged price | -0.68 (-1.94)** | ||||
| -0.33 (-2.04)** | |||||
| 0.65 (0.56) | |||||
| 0.08 (0.13) | |||||
| Worker | -0.16 (-3.02)*** | 0.07 (0.77) | -0.09 (-0.53) | -0.08 (-1.21) | |
| Size | -0.04 (-1.68)** | -0.03 (-0.81) | -0.01 (-0.16) | -0.14 (-4.49)*** | |
| Social Class 1 | -0.22 (-2.80)*** | -0.44 (-2.98)*** | -0.75 (-2.57)*** | -0.36 (-3.36)*** | |
| Social Class 2 | -0.06 (-0.94) | -0.40 (-3.44)*** | -0.40 (-1.74)** | 0.04 (0.49) | |
| Dummy of segments 1 | -0.66 (-4.46)*** | ||||
| Dummy of segments 2 | -0.19 (-2.02)** | ||||
| Constant | -0.81 (-4.97)*** | 29.92 (19.13)*** | -0.20 (-0.34) | 4.11 (21.00)*** | |
| -11.02 (-15.79)*** | 1.28 (3.05)*** | ||||
| 0.61 (2.25)*** | -5.03 (-3.76)*** | ||||
| 3.70 (4.92)*** | |||||
| MEASURES OF FIT | |||||
| Log likelihood | -5,340 | -3,614 | -4,119 | -4,254 | |
| # of purchases | 15,402 | 15,402 | 13,903 | 15,402 | 15,402 |
| # of segments(s) | S = 1 | S1 = 0.62 | S1 = 0.22 | S1 = 0.62 | |
| S2 = 0.26 | S2 = 0.12 | S2 = 0.26 | |||
| S3 = 0.12 | S3 = 0.60 | S3 = 0.12 | |||
| # of parameters | 7 | 13 | 15 | 70 | 11 |
| AIC | 10,694 | 7,254 | 8,268 | 10,188 | |
| HQ | 10,712 | 7,287 | 8,306 | 12,315 | |
| BIC | 10,815 | 7,478 | 8,524 | 24,707 | |
| CAIC | 10,756 | 7,368 | 8,398 | 17,449 | |
| Pseudo R2 | 0.36; 0.33; 0.38; 0.39; 0.39; 0.42; 0.43; 0.50; 0.43; 0.42; 0.41; 0.41 | ||||
Notes: 1. Models: Model 1. Logit model with state dependence, observed heterogeneity and without unobserved heterogeneity. Model 2. Latent class logit model with state dependence and observed heterogeneity. Model 3. Latent class logit model without state dependence and with observed heterogeneity and lagged price. Proposed Model. Panel model with state dependence and observed heterogeneity controlling unobserved heterogeneity using random effect. Model with two stages. 2. Table entries are the value coefficient with the t-statistic in parentheses, *p < 0.10, **p < 0.05 and ***p < 0.01. 3. Akaike Information Criterion (AIC = –2LogL + 2K); Hannan–Quinn (HQ = –2LogL + 2Kln(Ln(N)); Bayesian Information Criterion (BIC = –2LogL + Kln(N)); and Consistent Akaike Information Criterion (CAIC = –2LogL + K(ln(N) + 1)). Here, LogL is the value of the log-likelihood function for each model, K is the number of parameters estimated, and N is the sample size. 4. Pseudo-R2 = 1 – (LogL/ LogLC), where LogL is the value of the log-likelihood function at the optimum for the complete model and LogLC is the value of the log-likelihood function at the optimum for a restricted model with only a constant as the explanatory variable. 5. Variables used in the first step are the age of the main buyer of the household, age of the wife in the household, number of children of different ages and number of adults, dummies for sex, dummies indicating the geographic location of the household and dummies indicating the population size of the geographic location of the household, and household expenditure on the category during the analyzed period. Age squared and age cubed of the main buyer of the household, age squared of the household wife, and interactions between the sociodemographic variables.
The results for Model 2 show that there is unobserved heterogeneity in preferences and prices. Model 1 is a particular case of Model 2 where the number of segments is equal to one. If we compare the measure of fit between both models, we choose Model 2 in both categories because they have higher values of log-likelihood and smaller values of AIC, HQ, BIC and CAIC.
The control of unobserved effects is important for explaining brand choice behavior. For example, in the fine laundry detergent category, we find three consumer segments. In these segments, the state dependence variable, the price and the constant have similar t-statistics; this implies that state dependence and unobserved heterogeneity have a similar impact on brand choice. The state dependence variable losses some capacity to explain choice behavior once we control for unobserved heterogeneity in prices and preferences. The probability to choose today the same brand that in the previous choice is smaller when the probability to choose the brand is conditioned on unobserved heterogeneity. We get similar results in the non-fine laundry detergent category.
We find signs of overstated state dependence in Model 1 not only because the explanatory capability of the last variable is smaller in Model 2 but also because its size is smaller. With this focus, we specified Model 3 with the aim of distinguishing between true and spurious state dependence.
Therefore, Model 3 constitutes a simple test for distinguishing state dependence from heterogeneity and serial correlation. Chamberlain (1978) noted that the key distinction between heterogeneity and state dependence is the dynamic response to exogenous variables. The results indicate that the lagged price coefficients are statistically significant in segments one and two in both categories. Chamberlain’s test results suggest that there are intertemporal dependencies in the deterministic part of the utilities, and hence, choice dynamics. However, this test cannot split state dependence from heterogeneity because of the assumption of absence of correlation between the individual effect and the covariates. We then need to use another method to split state dependence and unobserved heterogeneity correctly in the case that correlation effects exist.
We must not forget in any case that: (i) the results from infraspecified models (without unobserved effects in Model 1) produce bias in the estimates; (ii) the results from the correctly specified model but with random unobserved heterogeneity (Model 2) can also produce bias when the state dependence variable is correlated with the individual effect.
To solve the correlation problem, we show the results from the two-stage estimation process in the proposed model. The determinants of unobserved heterogeneity in the reduced form model in the first estimation stage are: age of the main buyer in the household, age of the wife in the household, number of children of different ages and the number of adults, dummies for sex, dummies indicating the geographic location of the household and dummies indicating the population size of the geographic location of the household, and household expenditure on the category during the analyzed period. We assume that these variables are strictly exogenous to identify model parameters. Our aim is to use a rich specification to obtain a good model fit. We include all the sociodemographic determinants, along with the squares of the variables, and variable interactions. We need a good model fit because we will use the latent variable prediction as the dependent variable in the second stage. In this way, we avoid unobserved latent variables in the model specification. We fit a logit model in each time period, and the first-stage are not less than 0.70 (fine laundry detergent) and 0.33 (non-fine laundry detergent), representing a logit regression of the utility of the choice of brand on all exogenous variables.
We model a reduced logit model, where it is assumed that all the coefficients will be the same during the estimation period. If correlation is not present, then the coefficients will be the same in the reduced model as in the proposed model. If the coefficients are different, we can use this result as a sign of the correlation between the individual preference effects and the covariates. We test whether individual preference effects are uncorrelated with the covariates in the first stage. To do this, we run a Wald test between the model that we propose and the reduced logit model. The Wald statistic is equal to 0.4976. for 210 restrictions and significance at the 0.01 level is 156.432 for fine laundry detergent, and the statistic is 0.1947. The for 770 restrictions and significance at the 0.01 level is major than 156.432 for non-fine laundry detergent. We cannot reject the hypothesis of different coefficients. Therefore, we interpret these results as driven by correlated effects. We then use a method like that proposed to split state dependence and unobserved heterogeneity when correlation effects are presents to obtain bias-free parameters.
In the second stage of the proposed model, we obtain the parameter of interest. The state dependence variable is less significant than in the other nested models, and there are other brand choice determinants that have more important explanatory power in brand choice than the last variable. With respect to the remainder variables in the proposed model, we point out that price is the most significant variable. Brand intrinsic preference also has an important effect. In the fine laundry detergent category, we find three segments with different characteristics. The biggest is Segment 1. It is very price sensible and has no preference for buying SBs. Segment 2 is the second segment in size and is the lowest in price sensitivity because it has a preference for buying SBs that are cheaper. Segment 3 is the smallest segment and shows median price sensitivity, and the preference for SBs is similar to that for NBs. In the non-fine laundry detergent category, we find similar results. We find also three segments. The biggest one is Segment 1. It is the most price sensible segment and has no preference for buying SBs. Segment 2 is the second segment in size and is less price sensitive than Segment 1, but neither prefers to buy SBs. Only Segment 3 show high preferences toward SBs. This is the smallest segment and displays the lowest price sensitivity. These results imply that after controlling for unobserved effects in a correct way, the marketing-mix variables have a greater influence on households’ brand choice than state dependence.
To summarize, we can say that at the same time as model specification becomes richer, the state dependence influence on brand choice becomes less significant. In our opinion, these results clarify the way to correctly identify true state against spurious state dependence.
We cannot compare the performance of the proposed model with Model 2 using measures of fit because the former is estimated using maximum likelihood and we use generalized method of moments (GMM). However, in the following section, we compare the predictive ability of the traditional maximum likelihood estimation procedure and the estimation procedure proposed.
4.3 Predictive model performance
A final stage in the model selection procedure is the evaluation of the forecasting performance of one or more selected models. We may consider an ‘out of sample’ interpolation to evaluate the forecasting power of Model 2 because it is the richer traditional maximum likelihood specified model with the proposed model.
One indicator of this forecasting power is the root mean squared error (RMSE) in predictive choice market share. We assume that, on each choice occasion, the brand with the highest predicted choice probability is purchased. The predicted choice market share of each brand is then obtained by aggregating predicted choices across all purchase occasions.
\[R M S E = \sqrt {\frac {1}{N} \sum_ {i = 1} ^ {N} \left(P (y _ {i}) - \hat {P} (y _ {\hat {i}})\right) ^ {2}} = \sqrt {\frac {1}{N} \sum_ {i = 1} ^ {N} \left(e _ {i}\right) ^ {2}}\]
Another indicator of forecasting power, allowing comparison of the performance of different models, is the percentage of correct predictions for each model. This follows directly from a prediction–realization table, where the value can be interpreted as the hit rate. is the diagonal value in the prediction–realization table, the proportion of times that alternative j is correctly predicted. Based on simulation experiments, Veal and Zimmermann (1992) recommend the use of the measure suggested by McFadden, Puig and Krischner, which is given by F1; , where is the proportion of times alternative j is predicted.
The model with the lower value of RMSE and highest value of the hit rate and F1 may be viewed as the model that has the best forecasting performance. Measures of RMSE, the hit rate and F1 are reported in Table 3.
Table 3. Forecasting results for the out-of-sample in the fine and non-fine laundry detergent category
| RMSE1 | Hit rate2 | F13 | |
| Fine laundry detergent category | |||
| Model 2 | 0.3789 | 0.8564 | 0.6941 |
| Proposed Model | 0.3415 | 0.8833 | 0.7619 |
| Non-fine laundry detergent category | |||
| Model 2 | 0.5032 | 0.7467 | 0.4198 |
| Proposed Model | 0.3899 | 0.8479 | 0.2849 |
| Notes:Model 2. Latent class logit model with state dependence and observed heterogeneity.Proposed Model. Panel model with state dependence and observed heterogeneity controlling unobserved heterogeneity using random effect. Model with two stages.1. Root mean squared error (RMSE) in predictive market share. We assume that, on each choice occasion, the brand with the highest predicted choice probability is purchased. Predicted market share of each brand is then obtained by aggregating predicted choices across all purchase occasions (N).RMSE = $\sqrt{\frac{1}{N}\sum_{i=1}^{N}\left(P(y_i)-\hat{P}(y_i)\right)^2}=\sqrt{\frac{1}{N}\sum_{i=1}^{N}\left(e_i\right)^2}$ 2. These follow directly from a prediction–realization table, where the value $\sum_{j=1}^{2}pjj$ can be interpreted as the hit rate. $\sum_{j=1}^{2}p_{jj}-p_{.j}^2$ 3. Measure suggested by McFadden, Puig and Krischner, which is given by F1; $F1=\frac{\sum_{j=1}^{2}p_{jj}-p_{.j}^2}{1-\sum_{j=1}^{2}p_{.j}^2}$ . | |||
The proposed model performs better out of sample predictions in both categories. We find that the proposed model could be then a correct model for splitting state dependence and unobserved effects, and at the same time it could be a good alternative to make out of sample prediction after estimating the specifications.
5. Conclusions and Implications
The paper’s main aim is to distinguish state dependence from unobserved heterogeneity in the estimation context of a discrete choice model in a binary setting.
The results show evidence of true and spurious state dependence in the choice process when we control for observed heterogeneity and unobserved heterogeneity by methods commonly used in the brand choice literature (Keane 1997). Empirical tests have showed as to control unobserved heterogeneity limit the state dependence explanatory power in both laundry detergent categories used to fit the model.
Advance to split unobserved preference heterogeneity and state dependence do not fully avoid the problem of biased parameters (Abramson et al. 2000); biased parameters would imply wrong policies application by the marketing manager. We find that biased parameters can probably due to the correlation between state dependence and unobserved preference heterogeneity, and the correlation appears to be due to the way in which the state dependence variable is built.
We find that when using traditional maximum likelihood approaches, the individual effects are correlated with the covariates. This novelty of this analysis is that formalizing the correlation structure by imposing some assumptions, it is possible to consider this correlation during the estimation process. It allows derivation of a reduced form from which we obtain reduced form predictions that allow structural parameters to be derived in a simple way when individual effects are uncorrelated with the covariates.
On the other hand, because we have a limited number of observations for each household in consumer panel data, it is not possible to estimate consistently the unobserved household effects using maximum likelihood procedures, and the inconsistent unobserved household effects is translated to the parameter of interest. With the estimation procedure proposed here, we are able obtain consistent estimators.
When we control for unobserved effects in the way proposed by Chamberlain (1984), state dependence becomes a less significant variable. From this result, we conclude that when households make the brand choice decision by conditioning on the environmental effects, observed characteristics and unobserved characteristics without state dependence have an important influence on choice. This result could be the reason for finding mixed evidence about the order of consumer brand choice behavior in disaggregate and aggregate choice models. Because the influence of state dependence is small in a disaggregate setting, this effect could disappear when we model it in an aggregate context and evidence of zero-order brand choice process can be found.
From a managerial perspective, we conclude that there are some variables in the household’s information set that are not available in our sample, and these variables are more important than state dependence for explaining choice behavior. These unavailable variables can cause inertia. For example, the household can choose the most accessible brand to avoid wasting time looking for a cheaper brand because the time wasted cannot compensate for the reduction in price from the process of comparing brands or the lack of information concerning differences among brands. It is possible that the cost of accessing all household information is greater than the benefits in the case of the detergent category. In any case, these variables are individual specific and time invariant and have an important role in explaining brand choice behavior. Therefore, greater household knowledge is necessary for understanding households’ brand choice decisions.
With this comment, we would not like to say that the information contained in household panel data is not sufficient, rather the opposite. Household panel data generally provide richer information about sociodemographics. This can be very useful for predicting market shares by using the correct model that controls for the unobserved variables that are taken into account by the household but that are not available to the researcher. On the other hand, when we model in this way, we can obtain very significant information on the price sensitivity of the consumer segments. With this information, we can apply differentiated strategies to each consumer segment.
Retailers may think about their strategic decisions to hold a store brand portfolio. Store brands sell at a lower price than national brands with good quality levels, so the retailer could launch different store brands positioned with different prices aimed at different customer segments—for example, premium quality store brands that are not priced lower than national brands and are targeted at less price-sensitive consumers.
For example, if a retailer would like to increase the market share of his SB in the categories analyzed in the segment with the biggest size, one way would be to increase preferences for the SB. In both categories, the consumer in these segments does not buy the SB, and a small increase in the price has a big impact in the SB’s market share in this segment. Because SBs are the cheapest alternative, this result leads us to think that this segment is more concerned about quality than price.
As in any research, this investigation has certain limitations that must be considered. First, the proposed model is estimated using two frequently purchased categories. We have used categories where the SB’s market shares are very different and where market conditions and consumer behavior are different. To check the findings and provide more robust results, it would be interesting to apply the model to high-involvement categories and see if the results are maintained. Second, instead of binary choice options, we could further extend the model in future research to the multinomial case to model the level of brand alternatives.
Appendix.
Table A.1 Descriptive statistics
| Fine laundry detergent category | Non-fine laundry detergent category | |||
| SBs | NBs | SBs | NBs | |
| Dependent variable | ||||
| Purchase occasions | 1,638(37.52%) | 2,728(62.48%) | 5,752(19.14%) | 24,298(79.75%) |
| Explanatory variables | ||||
| Price (€) | 0.89(0.32) | 2.51(1.16) | 1.04(0.39) | 1.86(0.76) |
| Last | 0.68(0.46) | 0.10(0.30) | 0.61(0.48) | 0.07(0.26) |
| Size | 3.67(0.92) | 3.62(1.04) | 3.70(0.94) | 3.68(0.96) |
| Social class1 | 0.20(0.40) | 0.21(0.40) | 0.16(0.37) | 0.21(0.41) |
| Social class2 | 0.63(0.48) | 0.60(0.48) | 0.65(0.47) | 0.61(0.48) |
| Social class3 | 0.16(0.36) | 0.18(0.38) | 0.18(0.38) | 0.16(0.37) |
| Worker | 0.25(0.43) | 0.34(0.47) | 0.27(0.44) | 0.33(0.47) |
| Number of brands | ||||
| 23 | 38 | 39 | 77 | |
| Notes:1. Store brands (SBs) and national brands (NBs)2. Table entries in the first line is the mean value with the standard deviation in parentheses. | ||||
References
- Abramson, C., Andrews, R. L., Currim, M. S. & Jones, M. (2000). Parameter Bias from Unobserved Effects in the Multinomial Logit Model of Consumer Choice. Journal of Marketing Research, 37 (November), 410- 426.
- ACNielsen., (2005). The Power of Private Label.
- Ailawadi, K. L., Gedenk, M. & Neslin, S. A. (1999). Heterogeneity and Purchase Event Feedback in Choice Models: An empirical Analysis with Implications for Model Building. International Journal of Research in Marketing, 16, 177-198.
- Ailawadi, K. & Keller K. (2004). Understanding Retail Branding: Conceptual Insights and Research Priorities. Journal of Retailing, 80, 331-342.
- Andrews, R.L., Ainslie A., & Currim, I.S., 2002. An Empirical Comparison of Logit Choice Models with Discrete versus Continuous Representations of Heterogeneity. Journal of Marketing Research, 39, 479-488.
- Bass, F.M. & Wind, J. (1995). Introduction to the Special Issue: Empirical Generalizations in Maketing. Marketing Science, 14, G1-5.
- Bover, O. & Arellano, M. (1988). Estimating Dynamic Limited Dependent Variable Models from Panel Data. Investigaciones Económicas, 21, 141- 165.
- Bucklin, R. E. & Gupta, S. (1992). Brand Choice, Purchase Incidence, and Segmentation: An Integrated Modeling Approach. Journal of Marketing Research, 29 (May), 201-215.
- Chamberlain, G. (1978), “On the Use of Panel Data,” paper presented at the Social Science Research Council Conference on life-cycle aspects of employment and the labor market, Mt. Kisco, NY.
- Chamberlain, G. (1984). Panel Data. In Intrilligator (Ed.) Handbook of Econometrics 2 (pp. 1248-1318). North-Holland, Amsterdam: Griliches.
- Chintagunta, P., Jain. D.C. & Vilcassim, N. J. (1991). 'Investigating Heterogeneity in Brand Preferences in Logit Models for Panel Data. Journal ofMarketing Research, 28 (November), 417-428.
- Dhar, S., Hoch, J., 1997. Why Store Brand Penetration Varies by Retailer. Marketing Science 16 (3), 208-227.
- Elrond, T. & Keane, M.P. (1995). A Factor-Analytic Probit Model for Representing the Market Structure in Panel Data. Journal of Marketing Research, 32 (February), 1-16.
- Erdem, T. (1996). A Dynamic Analysis of Market Structure Using Panel Data. Marketing Science, 15 (4), 359-378.
- Erdem, T., & Sun, B. (2001).Testing for Choice Dynamics in Panel Data. Journal of Business & Economic Statistics, 19 (2), 142-152.
- Fader, P. S. & Lattin, J.M. (1993). Accounting for Heterogeneity and Nonstationarity in a Cross-Sectional Model of Consumer Purchase Behavior. Marketing Science, 12 (3), 304-317.
- Gönül, F. & Srinivasan, K. (1993). Modeling Multiple Sources of Heterogeneity in Multinomial Logit Models: Methodological and Managerial Issues. Marketing Science, 12 (3), 213-229.
- Guadagni, P. M. & Little, J. D. (1983). A Logit Model of Brand Choice Calibrated on Scanner Data. Marketing Science, 2 (3), 203-238.
- Gupta, S., & Chintagunta, P.K., 1994. On Using Demographic Variables to Determine Segment Membership in Logit Mixture Models. Journal of Marketing Research, 31, 128-136.
- Heckman, J.J. (1981). Heterogeneity and State Dependence. In Studies in Labor Markets (pp. 91-139). S. Rosen, Chicago: University of Chicago Press.
- Hsiao, C. (1986). Analysis of Panel Data. Cambridge, UK: Cambridge University Press.
- Jones, A. M. & Labeaga, J.M. (2003). Individual Heterogeneity and Censoring in Panel Data Estimates of Tobacco Expenditure. Journal of Applied Econometrics, 18, 157-177.
- Jones, M. J. & Landwehr, J. T. (1988). Removing Heterogeneity Bias from Logit Model Estimation. Marketing Science, 7 (Winter), 41-59.
- Kamakura, W. A. & Russell, G. J. (1989). A Probabilistic Choice Model for Market Segmentation and Elasticity Structure. Journal of Marketing Research, 26 (November), 379-390.
- Keane, M. P. (1997). Modeling Heterogeneity and State Dependence in Consumer Choice Behavior. Journal of Business and Economic Statistics, 15 (3), 310-327.
- Krishnamurthi, L. & Raj, S.P., (1988). A Model of Brand Choice and Purchase Quantity Price sensitivities. Marketing Science 7, 1-20.
- Lattin, J. M. (1987). A Model of Balanced Choice Behavior. Marketing Science, 6 (Winter), 48-65.
- Ortmeyer, G., Lattin, J.M. & Montgomery, D.B. (1991). Individual Differences in Response to Consumer Promotions. International Journal of Research in Marketing, 8, 169-186.
- Rossi, P. E. & Allenby, G.M. (1993). A Bayesian Approach to Estimating Household Parameters. Journal of Marketing Research, 30 (2), 171-82.
- Roy, R., Chintagunta, P.K. & Haldar, S. (1996). A Framework for Investigating Habits, “The Hand of the Past,” and Heterogeneity in Dynamic Brand Choice. Marketing Science, 15 (3), 280-299.
- Seetharaman. P.B. (2004). Modeling Multiple Sources of State Dependence in Random Utility Models: A Distributed Lag Approach. Marketing Science, 23 (2), 280-299.
- Uncles, M., Ehrenberg, A., & Hammond, K. (1995) Patterns of Buyer Behavior: Regularities, Models, and Extensions. Marketing Science, 14, G71-78.
- Varki, S. & Chintagunta, P. K. (2004). The Augmented Latent Class Model: Incorporating Additional Heterogeneity in the Latent Class Model for Panel Data. Journal of Marketing Research, 41 (May), 226-233.
- Veall, M.R., Zimmermann, K.F., 1992. Performance Measures from Prediction-Realization Tables. Economics Letters, 39, 129-134.
- Wedel, M. & Kamakura, W. (1998). Market Segmentation. Conceptual and Methodological Foundations. Boston: Kluwer Academic Publishers.