fedea
Fundación de Estudios de Economía Aplicada
Strategic behavior in regressions: an experimental by Javier Perote ** Juan Perote-Peña *** Marc Vorsatz Documento de Trabajo 2012-07
October 2012
* Universidad de Salamanca. ** Universidad de Zaragoza *** Universidad Nacional de Educación a Distancia and FEDEA.
Strategic behavior in regressions: an experimental study
Javier Perote∗ Juan Perote-Peña† Marc Vorsatz‡
October 15, 2012
Abstract
We study experimentally in the laboratory the situation when individuals have to report their private information (that is commonly known to be the sum of an observable and a random component) to a public authority that then makes inference about the true value hold by each of the individuals. It is assumed that individuals prefer this inferred or predicted value to be as close as possible to the their true value. Consistent with the theoretical literature, we show that the participants in our experiment misrepresent their private information more under the OLS than under the resistant line estimator (which extends the median voter theorem to the two–dimensional setting). Moreover, only the resistant line estimator is empirically unbiased and subjects earn significantly less if the OLS estimator is applied.
Keywords: Linear regression, robust estimation, laboratory experiment, resistant line, strategy-proofness.
JEL-Numbers: C10, C91, D70.
∗Corresponding author. Departmento de Economía, Universidad de Salamanca, Campus Miguel de Unamuno (Ed. FES), 37007 Salamanca, Spain. Email: perote@usal.es. Financial support from the Junta de Castilla y León, through the project SA218A11–1, is gratefully acknowledged
†Departmento de Análisis Económico, Universidad de Zaragoza, Gran Vía 2, 50005 Zaragoza, Spain. Email: jperote@unizar.es.
‡Departmento de Análisis Económico II, Universidad Nacional de Eduación a Distancia, Paseo Senda del Rey 11, 28040 Madrid and Fundación de Estudios de Economía Aplicada (FEDEA), Calle Jorge Juan 46, 28001 Madrid, Spain. Email: mvorsatz@cee.uned.es. Financial support from the Spanish Ministry of Education and Science, through the project ECO2009–07530, is gratefully acknowledged.
1 Introduction
Motivation Consider a simple linear regression model set up to estimate the values of a dependent variable conditional on the given values of an independent variable that represents the type of the individuals. Contrary to the literature in econometrics, we assume that the regression model not only helps to extract information about the underlying relationship between the two variables, but also that it is used to allocate resources among individuals in the future. For example, we are interested in problems such as the implementation of an income tax (see, for example, Saporiti 2009), the design of a cost sharing scheme (see, for example, Thomson 1983 and Sprumont 1991), or the construction of an incentive program based on estimated productivities (see, for example, Lazear 2000). The mentioned situations have in common that the dependent variable is unobservable —i.e, the agents’ willingness to pay for a service, their subjective valuation of public services, or the individual productivity or efort exerted at work—, and the regressions must rely on reported information.
In this setting, it might be the case that individuals have incentives to manipulate the regression output to their advantage if classical techniques like the OLS method are applied. An interesting and quite general case in which this occurs is when individuals are better of the closer the regression predictions (based on their reported information) are to their true private information. The following presents an example of this preference structure: consider a set of divisions within a big corporation that are asked to report their current expenditure that is private information and will not be revealed with certainty until the end of the year. The expenditure is a function of the number of workers in each division (or the capital invested) and some random efects. The divisions are asked to report their actual expenditure in order to design the optimal budget allocation among divisions for the next year. Some divisions that overspend might think that reporting the true expenditure could harm their longterm interests by inducing the managers to believe that their performance is below average and that they deserve to be “punished”. So, these divisions have an incentive to report lower valuations in order to get the regression line (the predicted expenditure given their investment level) closer to their true data. Similarly, divisions that underspend could fear that their above average performance relative to their investment might be interpreted as higher productivity and their next year funding could be reduced or their future targets be risen. These divisions will gain by exaggerating their true performance to bring the regression line closer to their true expenditures. We therefore assume that the individuals reporting the data always prefer to have the predicted value corresponding to their type as close as possible to their true private information. This kind of preferences over predicted values are called single–peaked preferences in the literature on voting and social choice theory.
The problem of regressions when individuals have single–peaked preferences call for the search of mechanisms that are strategy–proof; that is, we look for estimators that provide individuals with incentives to reveal their private information truthfully.1 These estimators can be obtained by using the properties of the median that have been proved to be strategy–proof in public goods allocation problems when individuals have single–peaked preferences on a single dimension (see, Moulin 1980). The extension of this “median voter” theorem to the two–dimensional context together with a whole family of strategy–proof estimators called “clockwise repeated median estimators” (CRM hereafter) can be found in Perote and Perote-Peña (2004). We will introduce these estimators formally in the next section.
Experiment While the theoretical results in Perote and Perote-Peña (2004) reveal that the class of CRM estimators outperforms the OLS estimator in terms of its manipulability when preferences a single–peaked, it is still an open question whether individuals take this adequately into account. To study this question, we run a laboratory experiment that is organized as follows: each subject in a groups of eight is assigned an observable variable x (called income) from the interval . The private information of individual i (called contribution) is equal to , where is normally distributed with mean zero and variance four. Individuals then report simultaneously their private information to the public authorities who then obtain predictions of the contributions from the data using either the OLS estimator (treatment OLS) or the resistant line estimator (treatment RL), one salient member of the class of all CRM estimators.
1Formally, the direct revelation mechanism is manipulable if there is some individual i and some strategy profile played by the other individuals such that not revealing the true preferences is a best response for individual i. The direct revelation mechanism is startegy– proof if and only if no individual can manipulate it; see Barber`a (2001).
In line with the hypotheses derived from the theoretical predictions, we find that the RL estimator outperforms the OLS estimator on all important dimensions. First, the degree of manipulation (the mean absolute diference between the reported and the true private information) is significantly greater under the OLS than under the RL estimator. In fact, the average manipulation amounts to 2.79 for the OLS and to 1.32 for the RL estimator. Interestingly, the diference between the two treatments is significant for all observable income levels. Also, we only find in treatment OLS that subjects with a higher income manipulate more than subjects with a lower income. This result suggests that subjects expected substantial manipulations when the OLS estimator is used (if they believed that the slope of the estimated line is close to one, the degree of the manipulation should not depend on the actual income). Second, it turns out that only the RL estimator is empirically unbiased; that is, the average estimated slope and the average estimated intercept are not statistically diferent from the true theoretical values. While we hypothesized that RL estimator performs relatively better than the OLS estimator on these terms, this even stronger finding clearly reveals the superiority of the RL estimator over the OLS estimator when the data is obtained from strategic individuals. Finally, a direct consequence of the former findings is that the social welfare is higher if the RL estimator is applied.
Remainder In the next section, the theoretical model is presented. Section 3 introduces the experimental design and derives the hypotheses. Afterwards, we present our experimental results. Finally, we conclude. Some estimation results and the translated instructions are relegated to the appendices.
2 Model
In this section, we introduce the formal model and the class of clockwise repeated median estimators (CRM estimators).
Consider a set of individuals indexed by i. There are two variables x and that take values in , so let D be the collection of all bi–dimensional data points . For all individuals is publicly observable, but is only known to individual i herself. To simplify the exposition, individuals are numbered in increasing order with respect to the vector , so that . Also, there is a known (albeit random) structure that connects x and . In particular, we follow the standard simple linear econometric model according to which . In this equation, u is a random vector distributed , where and I is the identity matrix, and is the vector of the parameters of the model.
This structure is supposed to be used to implement some policy depending on individual performances, which are measured in terms of the deviations of from the (estimated) conditional mean of variable y given x, . This is to say that individual i is better of the lower the distance is; that is, individual i is better of the closer the estimate is to the true value . Individuals have to report their private information to compute the estimates of the , and we denote by the value individual i reports when her true realization is .
Under the assumption of the simple linear econometric model, is the best linear predictor for provided that and are obtained by OLS from the true sample data. Obtaining such a performance measure needs to overcome two main problems. First, the traditional econometric estimation of requires the researcher to impose an assumption on the data generating process of . As explained before, we consider here the traditional simple linear regression model. Second, if is not observable and we base the estimation on the reported values , individuals may have incentives to report false information to improve their performance. For example, in this context the OLS estimator is the one that minimizes the sum of the squared residuals , where . This estimator is clearly manipulable because an individual has incentives to reduce the distance between her true and her estimated value by declaring higher (lower) values whenever the residual for the true sample is positive (negative).
In order to avoid the problem of strategic data manipulation, one can apply an alternative strategy–proof estimator from the class of all CRM estimators. An especially attractive member of this family is the well–known resistant line (RL) method (see, Tukey 1970), a simple regression method that is robust to the appearance of outliers and satisfies further important statistical properties (see, Johnstone and Velleman 1985). In order to introduce the resistant line method, we first have to specify a partition of the x−values into three groups. In particular, we have to choose two individuals such that and both l and are odd. The individuals l and r are used to divide the data set D into a left sample and a right sample . The slope estimate of the resistant line method is defined as the solution for to the following equation:
\[\underset {i \leq l} {\text { Median }} \{\tilde {y} _ {i} - \beta_ {2} x _ {i} \} = \underset {i > r} {\text { Median }} \{\tilde {y} _ {i} - \beta_ {2} x _ {i} \}.\]
This method amounts to finding graphically the line that has for both the left and the right sample the equal number of points above and below it. Since there is an odd number of observations in both samples, the regression line must pass through at least one observation of each group. There exist several algorithms to calculate the resistant line, but we shall use here the strategy of the clockwise angles technique, since it is quite simple and instrumental in defining all members of the CRM estimators class. According to it, we have to apply the following three–step approach.
1. For all individuals and all individuals , determine the angle that is obtained if observation is connected with observation by a straight line.2
(xi, y˜i)
2Observe that the angle of a vector that points from an observation in the bi–
2. Then, for all individuals , calculate the median angle of subject i as
3. Finally, calculate the directing angle an . The estimate of the slope is then just the tangent of the directing angle; that is, . Then, if individuals and define the directing angle so that the regression passes through the observations and , the estimate of intercept is the intercept corresponding to that line; that is,
\[\widehat {\beta} _ {1, R L} = \frac {x _ {j} \cdot \tilde {y} _ {j} - x _ {i} \cdot \tilde {y} _ {i}}{x _ {j} - x _ {i}}.\]
Note that the resistant line always exists and that it is unique under our conditions. A graphical example of this algorithm is given in the instructions of the experiments in the appendix.
Let us now provide an intuition about why the resistant line method is strategy–proof. Given any vector of reports such that individual declares her private value truthfully , suppose that the estimation process results in a situation in which the true value of individual i lies strictly above the resistant line; that is, . In this case, if individual i declares any other value , the resistant line will not be afected. Hence, she cannot gain by deviating in this direction. On the other hand, if she declares any value 2 the resistant line either does not move or shifts downwards, which implies that the prediction for individual i does not get better grows or remains the same). In fact, the only way for such an individual i to afect the resistant line is by reporting a y−value that jumps over the existing line and in case of shifting the regression line, it always takes it further away from the true value . Consequently, individual i cannot gain from these deviations either. An identical reasoning holds for all observations lying below the existing resistant line and those such that Consider finally the individuals dimensional space to another observation —remember that since , the second observation is always to the right of the first one— is simply the angle defined by the vector to the north (counter–clockwise).
that serve as support for the resistant line. They can clearly shift the resistant line in both directions, but since they get their most preferred outcomes, they cannot gain from misrepresenting their private information. Finally, since no individual has incentives to misreport her private value, we can conclude that the resistant line estimator is strategy–proof.
Despite the theoretical advantages of this method when strategic individuals provide the data, it can be argued that in practice agents may not be aware about its good strategic properties and are thus still going to manipulate the data. In what follows, we design an experiment to investigate this question.
3 Experiment
3.1 Setting
We applied a between–subjects design to see whether the RL estimator performs better than the OLS estimator in assessing the private information of individuals. We framed the experiment in the context of a tax declaration problem in order to help subjects to better understand the general environment. At the beginning of the game, every individual in a group of eight subjects gets assigned her income , which is observable to all participants (the exact incomes are known to the subjects, but we never revealed the actual mapping between subjects and incomes). Each subjects then privately observes her contribution , which is randomly drawn from the publicly known data generating process
\[y _ {i} = x _ {i} + u _ {i}, \text { where } u _ {i} \sim N (0, 4).\]
Observe that and . The participants are then asked to simultaneously and independently report their contribution. The revealed contribution or report of subject i can be any rational number from the interval [0,24]. Given a vector of reports , inference about the true contribution is made using either the OLS or the RL estimator. The exact procedure is known to the subjects.
1. Treatment 1: OLS
If the OLS estimator is used,
\[\widehat {\beta} _ {2, O L S} = \frac {\sum_ {i = 1} ^ {n} (x _ {i} - \bar {x}) (\tilde {y} _ {i} - \bar {\tilde {y}})}{\sum_ {i = 1} ^ {n} (x _ {i} - \bar {x}) ^ {2}}\]
and
\[\widehat {\beta} _ {1, O L S} = \bar {\tilde {y}} - \hat {\beta} _ {2, O L S} \bar {x},\]
where and
2. Treatment 2: RL
If the RL estimator is used, the following procedure is applied. First, we define the sets and containing the first three and the last three subjects, respectively, and compute the nine angles formed if each report from the members of L is connected with each report from the members of R. The median angle of subject is then . The directing angle is then given by . Finally, the estimate of the slope is obtained as the tangent of the directing angle,
\[\widehat {\beta} _ {2, R L} = \tan (D A),\]
and the estimate of the intercept is the one that corresponds to the two observations of the sample that are defining the directing angle; that is, if the regression passes through the observations and , then
\[\widehat {\beta} _ {1, R L} = \frac {x _ {j} \cdot \tilde {y} _ {j} - x _ {i} \cdot \tilde {y} _ {i}}{x _ {j} - x _ {i}}.\]
Once the estimator is obtained from the reports, the fitted contribution for subject i is calculated as
\[\widehat {y} _ {i} = \widehat {\beta} _ {1} + \widehat {\beta} _ {2} \cdot x _ {i}.\]
Finally, subjects receive their payof as a function of their true and their fitted contribution:
\[\pi_ {i} (y _ {i}, \widehat {y} _ {i}) = \max \{5 - | y _ {i} - \widehat {y} _ {i} |, 0 \}.\]
3.2 Procedures
We conducted the experiment, which was programmed within the z–Tree toolbox (see, Fischbacher 2007), in the Laboratory for Research in Social and Economic Behavior (LINEEX), which is hosted at the University of Valencia. For each treatment, we organized one session with 8 subjects and another one with 40 subjects. Hence, in total, 96 undergraduates from various disciplines participated in one of the experimental sessions.
Before the start of a session, participants privately read the instructions that included a detailed example of the estimation technique (see, the appendix). The subjects were then able to test their understanding of the instructions in six practice rounds that did not afect their final payof. After the completion of the practice rounds, the participants had to answer several control questions. The software only started once all participants answered all control question correctly.
Participants were then randomly assigned into groups of eight. Their identities were never revealed. To ensure that the data is truly independent across groups, the participants were also informed that they would only play against subjects from the same group and that the group assignment would not change during the experiment. Within each group, the game was played 48 times. The 48 rounds were divided into 6 blocks of 8 rounds. Subjects were assigned incomes (types) in such a way that within each block, every subject had once an income of 2, once an income of 4, and so forth. To maximize the comparability of the treatments, one series of error terms of size 8 (subjects per group) × 6 (number of groups) × 48 (number of periods) was drawn. This series of error terms was then used in both treatments.
In each round, after having learned their income and their contribution, subjects submitted their reports. Given the vector of reports, the fitted contributions were determined and the participants received their payofs. At the end of round, the subjects were presented a summary screen that included their income, their contribution, their report, their fitted contribution, the diference between their true and their fitted contribution, and their payof. The same information was also graphically presented together with the estimated line
\[\widehat {y} = \widehat {\beta} _ {1} + \widehat {\beta} _ {2} \cdot x\]
and the 95 % confidence interval of the contribution (see, the instructions in the appendix). Observe that the participants never received information on the true or reported contribution of their co–players.
Participants earned experimental currency units (ECUs) during the experiment that were converted into Euros at a known exchange rate at the end of the experiment. Payment took place privately and the students had to leave the laboratory immediately once paid. The average payof was 17.03 Euros in treatment OLS and 17.36 Euros in treatment RL. A session lasted on average approximately 105 minutes.
3.3 Hypotheses
We now derive the experimental hypotheses that follow straightforwardly from the theoretical analysis in Section 2. Most importantly, since the RL estimator is strategy–proof and the OLS estimator is manipulable, we expect that the average absolute diference between the reported and the true contribution is larger in treatment OLS than in treatment RL. Observe that we do not expect the diference to be zero in treatment RL as predicted by the theoretical model, since it is very likely that many subjects will not realize that it is in their best interest to report their true contribution.
Hypothesis 1 (Manipulations): The average absolute diference between the reported and the true contribution (the degree of manipulation) is larger in treatment OLS than in treatment RL.
It is easy to see that if all subjects report their private information truthfully, then both estimators are unbiased and the expected payof of the subjects is maximal for the OLS. However, if subjects manipulate more under the OLS estimator than under the RL estimator, as it has been predicted in Hypothesis 1, the strategic interaction leads to a worse outcome in treatment OLS. Consequently, we expect the average estimate of the OLS estimator to be further away from the true parameters of the true underlying data generating process than the average estimate of the RL estimator
Hypothesis 2 (Biasedness): The OLS estimator is more biased than the RL estimator
Following exactly the same line of argumentation as above, under Hypothesis 1 the average payof should be higher under the RL estimator than under the OLS estimator. This hypothesis highlights that there are negative welfare efects under the OLS estimator if the data is revealed strategically and individuals have single–peaked preferences.
Hypothesis 3 (Payoffs): The average payof is higher in treatment RL than in treatment OLS.
4 Results
This section is divided into three parts. We study first how the subjects manipulate the estimators (Hypothesis 1). Afterwards, we analyze if, as a consequence of individual behavior, the estimators are biased (Hypothesis 2). Finally, we analyze welfare (Hypothesis 3).
In our statistical analysis, we proceed as follows. First, we calculate for each group the averages of the variables of interest over all rounds: the diference between the true and the reported contribution, the estimated slope and intercept, and the payof. This results in six truly independent observations (one per group). We then apply Wilcoxon signed–rank tests for within treatment comparisons and Mann Whitney U tests for between treatment comparisons.
4.1 Manipulations
Our first hypothesis states that subject manipulate more if the OLS estimator is applied. To see whether this is true, we plot in Figure 1 the average absolute diference between the true and the reported contributions. The values are averaged over three rounds and all group.
Figure 1: Average manipulation (absolute value of the diference between the true and the reported contribution) for the OLS treatment (bullets) and the RL treatment (circles) over rounds (3–round averages).

The figure shows that subjects deviate more from their true contribution under the OLS than under the RL estimator in all periods of the experiment. In fact, the average manipulation in treatment OLS is about 3.5 ECU in the beginning of the experiment and declines slightly over time to values between 2.5 and 3 ECU. On the other hand, in treatment RL, the average deviation starts at about 1.5 ECU, where it roughly remains until the end of the experiment. Overall, the average manipulation is 2.79 ECU under the OLS but only 1.32 ECU under the RL estimator. A Mann Whitney U test establishes that this diference is significant —the one–sided p−value is equal to 0.0039— and therefore, we can conclude that the subjects’ reports are further away from their true contribution under the OLS than under the RL estimator.
Table 1: Average manipulation (absolute value of the diference between the true and the reported contribution) for the OLS and the RL treatment by income. In brackets, the two– sided p−values of the corresponding Mann Whitney U tests at the group level.
| Treatment | Income in ECU | |||||||
| 2 | 4 | 6 | 8 | 10 | 12 | 14 | 16 | |
| OLS | 1.92[0.0250] | 2.28[0.0160] | 2.11[0.0250] | 2.80[0.0066] | 2.53[0.0250] | 2.93[0.0039] | 3.48[0.0039] | 4.23[0.0039] |
| RL | 1.29 | 1.25 | 1.18 | 1.42 | 1.26 | 1.46 | 1.21 | 1.48 |
One question that emerges at this point is whether the RL estimator works better in terms of manipulations than the OLS estimator for all income levels. The relevant data is presented in Table 1. It can indeed be seen that subjects deviate more from their true contributions level under the OLS than under the RL estimator independently of their income level. This insight is fully supported by the statistical analysis (see the p−values in the table).
Table 1 also reveals that the degree of manipulation increases with the income in treatment OLS. For example, while subjects with an income between 2 and 6 ECU deviate from their true contribution by about 2 ECU, subjects with an income of 14 ECU manipulate on average by 3.48 ECU, and subjects with an income of 16 ECU deviate even more than 4 units. Some part of this trend can certainly be attributed to the fact that the reports in our experiment are restricted to be non–negative, however this limitation only afects subjects with the lowest incomes and it cannot account for the fact that the subjects with an income of 16 ECU deviate significantly more from their true contribution than all other income groups (the two–sided p−values of the corresponding Wilcoxon signed rank tests are between 0.0277 and 0.0464).
The picture one gets in treatment RL is very diferent. The group that manipulates least is not the one with the lowest income but the one that contains the subjects with an income of 6 ECU (they deviate on average by 1.18 ECU). Subjects with an income of 16 ECU still manipulate more than all other income groups, but the diference is now rather negligible: these subjects deviate on average by only 0.40 ECU more than the group that manipulates least. From a statistical point of view, it turns out that only 3 of the 28 possible pairwise comparisons are significant at the five percent level: the comparison between subjects with an income of 16 ECU and those with an income of 6, 4, and 12 ECU, respectively. Consequently, we summarize our results so far as follows.
Result 1: Subjects manipulate more in treatment OLS than in treatment RL. This is true for all periods and all income levels. The degree of manipulation increases with the income in treatment OLS but not in treatment RL.
4.2 Biasedness
We have seen in the first part of our analysis that subjects indeed manipulate more if the OLS estimator is used. Next, we are going to study the consequences of these manipulations for the properties of the estimators.
In principle, there is the possibility that the larger deviations from the true contributions cancel out in such a way that the OLS estimator remains unbiased (the estimated parameters are equal to the true underlying process), however for this to happen the diferent manipulations must exactly ofset each other. Since this seems highly unlikely, our second hypothesis states that only the RL estimator is unbiased. In order to evaluate the hypothesis, Table 2 presents the average fitted intercepts and the average fitted slopes.3. Remember that, according the true process, the intercept equals zero and the slope equals 1.
Table 2: Average estimated intercept and slope for the OLS and the RL treatment. In parenthesis, the two–sided p−values of the Wilcoxon signed–rank tests at the group level that analyze whether the estimates are equal to the true underlying process. In brackets, the two– sided p−values of the Mann Whitney U tests at the group level that analyze the equality of the estimates across treatments.
| Estimates | Treatment | ||
| OLS | RL | ||
| Intercept | 1.4467(0.0464) | [0.0250] | -0.0966(0.9165) |
| Slope | 0.8019(0.0277) | [0.0104] | 0.9606(0.2489) |
It can be seen from the table that the OLS estimator is highly biased: the average fitted intercepts is significantly larger than zero and, most importantly, the average fitted slope is significantly smaller than one. On the other hand, we cannot reject the hypothesis that the RL is unbiased at the five percent significance level. Indeed, both the average fitted intercept and the average fitted slope are very close to the true values.
3Figures 2 and 3 in the appendix present the average estimates at the group level. Observe that the intercept and the slope of the presented graphs correspond to the independent observations of our statistical analysis
Result 2: Only the RL estimator is empirically unbiased.
4.3 Payofs
By definition, the OLS estimator minimizes the sum of the squared residuals, which is not the case for the RL estimator. Hence, if subjects always reported their contributions truthfully, the final payofs would necessarily be higher in treatment OLS than in treatment RL (because of the single–peakedness of the utility function). However, since we have shown in Result 2 that only the RL estimator is unbiased, it is a priori not clear in which treatment subjects fare better.
If we compare the relevant numbers, it turns out that subjects earn on average 2.91 ECU per period in treatment OLS and 3.26 ECU in treatment RL. Since this diference turns out to be significant (the two–sided p−value of the corresponding Mann Whitney U test is 0.0104), there is statistical evidence that even though the OLS estimator minimizes the sum of the squared residuals, the RL estimator leads to fitted values that are closer to the true contribution levels. Hence, the loss in the eficiency of the estimator is more than compensated by strategy–proofness.
Finally, and similar to our analysis with respect to the manipulations, we are going to study whether the payof of the subjects depends on the income level. The corresponding data is presented in the next table.
Table 3: Average payof for the OLS and the RL treatment by income. In brackets, the two–sided p−values of the corresponding Mann Whitney U tests at the group level.
| Treatment | Income in ECU | |||||||
| 2 | 4 | 6 | 8 | 10 | 12 | 14 | 16 | |
| OLS | 2.92[0.0782] | 3.00[0.2623] | 3.26[0.1630] | 2.95[0.2002] | 3.23[0.7488] | 2.89[0.6310] | 2.61[0.0374] | 2.41[0.0104] |
| RL | 3.27 | 3.32 | 3.70 | 3.70 | 3.25 | 3.14 | 3.35 | 3.28 |
It can be seen from Table 3 that for all income levels, subjects earn more in treatment RL than in treatment OLS. Yet, the diference is only significant at the five percent level if the income is 14 or 16 ECU, the two income levels where subjects manipulated most and earn least in treatment OLS.
Result 3: Subjects earn more in treatment RL than in treatment OLS; that is, the RL estimator leads to values that are closer to the true underlying contributions than the OLS estimator.
5 Conclusion
In this paper, we have designed a laboratory experiment in order to study the performance of the OLS and the resistant line estimator when the dependent variable is unobservable and the corresponding data is gathered from the reports of strategic individuals. It is well known from the theoretical literature that if preferences are single–peaked (that is, the individuals prefer their estimated value to be as close as possible to their private information), then individuals have incentives to misrepresent their private information under the OLS but not under the resistant line estimator. Our experimental results fully confirm the superiority of the resistant line estimator for this case. In fact, we find that (1) subjects deviate more from their true private information under the OLS than under the RL estimator, (2) only the RL estimator is empirically unbiased, and (3) subjects earn significantly more under RL than under the OLS estimator. Our results therefore highlight that the OLS estimation procedure should be used with care whenever the dependent variable is obtained from individual reports and the payof of the individuals depends on the estimation results. In these cases, alternative strategy–poof estimation techniques should be investigated and implemented.
References
- [1] Barber`a, S. (2001). An introduction to strategy-proof social choice functions. Social Choice and Welfare 18: 619–653.
- [2] Fischbacher, U. (2007). Z-Tree — Zurich toolbox for readymade economic experiments. Experimental Economics 10: 171–178.
- [3] Johnstone, I. and Velleman, P. (1985). The resistant line and related regression methods. Journal of the American Statistical Association 80: 1041– 1054.
- [4] Lazear, E. (2000). Performance pay and productivity. American Economic Review 90: 1346–1361.
- [5] Moulin, H. (1980). On strategy-proofness and single-peakedness. Public choice 35: 437-455.
- [6] Perote, J. Perote-Peña, J. (2004). Strategy-proof estimators for simple regression. Mathematical Social Sciences 47: 153–176.
- [7] Saporiti, A. (2009). Strategy-proofness and single-crossing. Theoretical Economics 4: 127–163.
- [8] Sprumont, Y. (1991). The division problem with single-peaked preferences: a characterization of the uniform allocation rule. Econometrica 59: 509– 519.
- [9] Thomson, W. (1983). Problems of fair division and the egalitarian principle. Journal of Economic Theory 31: 211-226.
- [10] Tukey, J. (1970). Exploratory data analysis. Adison-Wesley, Reading, MA (Limited Preliminary Edition).
Estimated Lines





Figure 2: The average fitted regression line (in red) and the true underlying process (in black) in the OLS treatment for each of the six groups.






Figure 3: The average fitted regression line (in red) and the true underlying process (in black) in the RL treatment for each of the six groups.

Instructions OLS (Translated from Spanish)
This experiment explores the design of an income tax system. You will be assigned to a group of 8 subjects that remains constant during the 48 rounds that the experiment lasts. In every round, you may earn a quantity measured in ECU (experimental currency units) that will be converted in euros at the end of the experiment at the rate
\[1 0 \mathrm{ECU} = 1 \mathrm{Euro}.\]
1. In every round, you will be assigned an income R and a contribution C. The participants from your group have diferent incomes of the following quantities: {2, 4, 6, 8, 10, 12, 14, 16}. The contribution of every group member depends on the income and will be randomly drawn from the following process:
\[C = R + e,\]
where e is a normally distributed random variable with mean zero and variance four. This means that if your income is R, then your contribution C will be in the interval , although with a small probability of 5% it may be outside this interval. Note that all participants from your group know the incomes but NOT the contributions of the other co-players; that is, every participant only knows her own contribution.
2. The only decision you have to take each period is to report a contribution. The reported contribution can be a (rational) number between 0 and 24.
3. Given the reported contributions of all group members, an estimation of the parameters of an income tax system will be computed: the intercept (lump-sum) and the slope (income percentage). This computation will be based on a simple rule that will be explained below in the section “estimation method”.
4. Given the estimates for the intercept and the slope, an estimated contribution will be computed for every subject in the following way:
\[C ^ {*} = i n t e r c e p t + s l o p e \times R.\]
5. Each round, the payof you receive will be the maximum of zero and
\[5 - | C - C ^ {*} |.\]
Consequently, your monetary benefits from the experiment will be the higher the closer the estimated contribution is to your true contribution.
Next, we display a figure as an illustration of those you will find throughout the experiment.

The straight line in green represents the estimated contributions calculated from the reported contributions of your group members in Period 1 Your income (R) is 10 Your contribution (C) is 15.0 (red point) Your reported contribution is 16.0 (black point) Your estimated contribution is 11.0 (yellow point) The error is 4.0 (blue line)

In the graph on the left hand side of the figure, the red lines capture the bands where the contributions of the eight subjects of your group should be placed with 95% probability. Your income (R) is 10, your contribution (the red point) is 15, and your reported contribution (the back point) is 16. The green line indicates the estimated contributions that are obtained from the reported contributions of all group members. In particular, the yellow point represents your own estimated contribution given the estimated income tax system.
On the right hand side of the figure, you find the values of the main variables, which will be collected in a table as the experiments progresses. Every period it is displayed the value of your income, your contribution (the red point), your reported contribution (the black point), your estimated contribution (the yellow point), the diference between your true and your estimated contribution (the blue line segment) and your payof from the period.
Estimation method
In every period, the estimation of the income tax line requires the estimation of both the “intercept” and the “slope” parameters given the known values of the income R and the reported contributions of the eight group members. Hereafter, we show an example which explains graphically the procedure to obtain these estimates assuming that reported contributions are those in the next picture below.

Given these observations, the estimated line (the green line in the figure below on the left hand side) will be the one that minimizes the sum of the squared vertical distances (errors) between the reported contributions and those of the estimated line. Note the sum the errors above (the blue lines) and below (the red lines) the estimated line are exactly the same. In the example, it is assumed that the reported contributions are the truly assigned contributions. The estimated contributions are the values of the contributions for every income level on the estimated line (the green line). The final payofs of the period are computed as 5 minus the distance between the true contribution and the estimated contributed.

| R | C | rev. C | est. C | Payoff |
| 2 | 4.4 | 4.4 | 5.0 | 4.4 |
| 4 | 2.3 | 2.3 | 6.3 | 1.0 |
| 6 | 10.8 | 10.8 | 7.6 | 1.8 |
| 8 | 6.5 | 6.5 | 8.9 | 2.6 |
| 10 | 15 | 15 | 10.3 | 0.3 |
| 12 | 15.8 | 15.8 | 11.6 | 0.8 |
| 14 | 12.1 | 12.1 | 12.9 | 0.7 |
| 16 | 10 | 10 | 14.3 | 0.7 |
We are now going to illustrate the impact of your reported contribution on payofs. Starting with the numbers in the table above, what would have happened if the subject with income 16 and the contribution 10 had reported a contribution of 2.3 (red point in the figure below)? You can observe the impact of such decision on the estimated line, which would have changed from the “light green” to the “dark green” line, in the plot below. The table on the right highlights the efects on the payofs for all subjects in the group. It is clear that the subject that changed her reported contribution would increase its payof from 0.7 to 3.9.

Finally, what would have happened if the subject with income 6 had reported 15.8 instead of her true contribution, 10.8? The following figure and
| R | C | rev. C | est. C | Payoff |
| 2 | 4.4 | 4.4 | 6.2 | 3.2 |
| 4 | 2.3 | 2.3 | 6.9 | 0.4 |
| 6 | 10.8 | 10.8 | 7.6 | 1.8 |
| 8 | 6.5 | 6.5 | 8.3 | 3.2 |
| 10 | 15 | 15 | 9 | 0 |
| 12 | 15.8 | 15.8 | 9.7 | 0 |
| 14 | 12.1 | 12.1 | 10.4 | 3.3 |
| 16 | 10 | 2.3 | 11.1 | 3.9 |
table illustrate this case. The subject with income 6 would increase its payof from 1.8 (her payof for the initial case where all subjects report their true contribution) to 2.7.

| R | C | rev. C | est. C | Payoff |
| 2 | 4.4 | 4.4 | 6.2 | 3.2 |
| 4 | 2.3 | 2.3 | 7.4 | 0 |
| 6 | 10.8 | 15.8 | 8.5 | 2.7 |
| 8 | 6.5 | 6.5 | 9.7 | 1.8 |
| 10 | 15 | 15 | 10.8 | 0.8 |
| 12 | 15.8 | 15.8 | 12 | 1.2 |
| 14 | 12.1 | 12.1 | 13.1 | 4 |
| 16 | 10 | 10 | 14.3 | 0.7 |
Now you will be able to practice in your computer with similar examples during six diferent periods. The payofs of these rounds will not afect your final payofs. Once you finish these examples and after filling out a brief questionnaire, the experiment will start. Remember that the experiment lasts 48 periods and you will play all of them within the same group composition.
Instructions RL (Translated from Spanish)
This experiment explores the design of an income tax system. You will be assigned to a group of 8 subjects that remains constant during the 48 rounds that the experiment lasts. In every round, you may earn a quantity measured in ECU (experimental currency units) that will be converted in euros at the end of the experiment at the rate
\[1 0 \mathrm{ECU} = 1 \mathrm{Euro}.\]
1. In every round, you will be assigned an income R and a contribution C. The participants from your group have diferent incomes of the following quantities: {2, 4, 6, 8, 10, 12, 14, 16}. The contribution of every group member depends on the income and will be randomly drawn from the following process:
\[C = R + e,\]
where e is a normally distributed random variable with mean zero and variance four. This means that if your income is R, then your contribution C will be in the interval , although with a small probability of 5% it may be outside this interval. Note that all participants from your group know the incomes but NOT the contributions of the other co-players; that is, every participant only knows her own contribution.
2. The only decision you have to take each period is to report a contribution. The reported contribution can be a (rational) number between 0 and 24.
3. Given the reported contributions of all group members, an estimation of the parameters of an income tax system will be computed: the intercept (lump-sum) and the slope (income percentage). This computation will be based on a simple rule that will be explained below in the section “estimation method”.
4. Given the estimates for the intercept and the slope, an estimated contribution will be computed for every subject in the following way:
\[C ^ {*} = i n t e r c e p t + s l o p e \times R.\]
5. Each round, the payof you receive will be the maximum of zero and
\[5 - | C - C ^ {*} |.\]
Consequently, your monetary benefits from the experiment will be the higher the closer the estimated contribution is to your true contribution.
Next, we display a figure as an illustration of those you will find throughout the experiment.

The straight line in green represents the estimated contributions calculated from the reported contributions of your group members in Period 1 Your income (R) is 10 Your contribution (C) is 15.0 (red point) Your reported contribution is 16.0 (black point) Your estimated contribution is 11.0 (yellow point) The error is 4.0 (blue line)

In the graph on the left hand side of the figure, the red lines capture the bands where the contributions of the eight subjects of your group should be placed with 95% probability. Your income (R) is 10, your contribution (the red point) is 15, and your reported contribution (the back point) is 16. The green line indicates the estimated contributions that are obtained from the reported contributions of all group members. In particular, the yellow point represents your own estimated contribution given the estimated income tax system.
On the right hand side of the figure, you find the values of the main variables, which will be collected in a table as the experiments progresses. Every period it is displayed the value of your income, your contribution (the red point), your reported contribution (the black point), your estimated contribution (the yellow point), the diference between your true and your estimated contribution (the blue line segment) and your payof from the period.
Estimation method
In every period, the estimation of the income tax line = intercept+slope×R requires the estimation of both the “intercept” and the “slope” parameters given the known values of the income R and the reported contributions of the eight group members. Hereafter, we show an example which explains graphically the procedure to obtain these estimates assuming that reported contributions are those in the next picture below.
Given these observations, the estimated line is obtained as follows.

1. Take the subject with income 2 and trace the vectors that pass through her reported contribution and those of the subjects with incomes 12, 14 y 16. From these three lines, choose the median or central one (the red line).

2. Take the subject with income 4 and trace the vectors that pass through her reported contribution and those of the subjects with incomes 12, 14 . From these three lines, choose the median or central one (the red line).

3. Take the subject with income 6 and trace the vectors that pass through her reported contribution and those of the subjects with incomes 12, 14 . From these three lines, choose the median or central one (the red line).
4. Now take the three median vectors chosen in the last three steps (associated with the observations of the subjects , respectively) and choose the median (central) one of these thee vectors as the estimated line.

Consequently, the estimated line always passes through two of the reported observations: those with the median reported contribution of the subjects with the three lowest (2, 4 an 6) and largest (12, 14 and 16) incomes. Furthermore observations of subjects with income 8 and 10 are always discarded in the procedure.

Given the estimated line in the figure above, the initial payof of 5 ECU for every participant will be reduced by the vertical distance between the true contribution C and the one corresponding to the estimated line. Let us assume that all subjects reported their true assigned contributions except for the subject with income 16, whose true contribution is 12.1, instead of 10, which is what she reported. In this case, payofs are 5 ECU minus the vertical distances from their reported values to the estimated ones (depicted in blue).

| R | C | rev. C | est. C | Payoff |
| 2 | 4.4 | 4.4 | 4.4 | 5 |
| 4 | 2.3 | 2.3 | 5.7 | 1.6 |
| 6 | 10.8 | 10.8 | 6.9 | 1.1 |
| 8 | 6.5 | 6.5 | 8.3 | 3.3 |
| 10 | 15 | 15 | 9.5 | 0 |
| 12 | 15.8 | 15.8 | 10.8 | 0 |
| 14 | 12.1 | 12.1 | 12.1 | 5 |
| 16 | 12.1 | 10 | 13.4 | 3.7 |
Note that if the subject with income 16 had reported her true contribution 12.1 or whichever other value less than13.3, she would have obtained the same payof since it would have not changed the estimated line (given the same values for all other subjects). Still, if he had reported a higher contribution than 13.3 (estimated contribution for all the true contributions), for example 15, the estimated line and her expected payof (and that of all other participants) would have changed. Finally, we present a figure and a table with the estimated line and corresponding payof in this case.

| R | C | rev. C | est. C | Payoff |
| 2 | 4.4 | 4.4 | 4.4 | 5 |
| 4 | 2.3 | 2.3 | 5.9 | 1.4 |
| 6 | 10.8 | 10.8 | 7.4 | 1.6 |
| 8 | 6.5 | 6.5 | 8.9 | 2.6 |
| 10 | 15 | 15 | 10.5 | 0.6 |
| 12 | 15.8 | 15.8 | 12 | 1.1 |
| 14 | 12.1 | 12.1 | 13.5 | 3.7 |
| 16 | 12.1 | 15 | 15 | 2.1 |
Now you will be able to practice in your computer with similar examples during six diferent periods. The payofs of these rounds will not afect your final payofs. Once you finish these examples and after filling out a brief questionnaire, the experiment will start. Remember that the experiment lasts 48 periods and you will play all of them within the same group composition.
References
- 2012-07: “Strategic behavior in regressions: an experimental”, Javier Perote, Juan Perote-Peña y Marc Vorsatz.
References
- 2012-06: “Access pricing, infrastructure investment and intermodal competition”, Ginés de Rus y M. Pilar Socorro.
References
- 2012-05: “Trade-offs between environmental regulation and market competition: airlines, emission trading systems and entry deterrence”, Cristina Barbot, Ofelia Betancor, M. Pilar Socorro y M. Fernanda Viecens.
References
- 2012-04: “Labor Income and the Design of Default Portfolios in Mandatory Pension Systems: An Application to Chile”, A. Sánchez Martín, S. Jiménez Martín, D. Robalino y F. Todeschini.
References
- 2012-03: “Spain 2011 Pension Reform”, J. Ignacio Conde-Ruiz y Clara I. Gonzalez.
References
- 2012-02: “Study Time and Scholarly Achievement in PISA”, Zöe Kuehn y Pedro Landeras.
References
- 2012-01: “Reforming an Insider-Outsider Labor Market: The Spanish Experience”, Samuel Bentolila, Juan J. Dolado y Juan F. Jimeno.
References
- 2011-13: “Infrastructure investment and incentives with supranational funding”, Ginés de Rus y M. Pilar Socorro.
References
- 2011-12: “The BCA of HSR. Should the Government Invest in High Speed Rail Infrastructure?”, Ginés de Rus.
References
- 2011-11: “La rentabilidad privada y fiscal de la educación en España y sus regiones”, Angel de la Fuente y Juan Francisco Jimeno.
References
- 2011-10: “Tradable Immigration Quotas”, Jesús Fernández-Huertas Moraga y Hillel Rapoport.
References
- 2011-09: “The Effects of Employment Uncertainty and Wealth Shocks on the Labor Supply and Claiming Behavior of Older American Workers”, Hugo Benítez-Silva, J. Ignacio García-Pérez y Sergi Jiménez-Martín.
References
- 2011-08: “The Effect of Public Sector Employment on Women’s Labour Martket Outcomes”, Brindusa Anghel, Sara de la Rica y Juan J. Dolado.
References
- 2011-07: “The peer group effect and the optimality properties of head and income taxes”, Francisco Martínez-Mora.
References
- 2011-06: “Public Preferences for Climate Change Policies: Evidence from Spain”, Michael Hanemann, Xavier Labandeira y María L. Loureiro.
References
- 2011-05: “A Matter of Weight? Hours of Work of Married Men and Women and Their Relative Physical Attractiveness”, Sonia Oreffice y Climent Quintana-Domeque.
References
- 2011-04: “Multilateral Resistance to Migration”, Simone Bertoli y Jesús Fernández-Huertas Moraga.
References
- 2011-03: “On the Utility Representation of Asymmetric Single-Peaked Preferences”, Francisco Martínez Mora y M. Socorro Puy.
References
- 2011-02: “Strategic Behaviour of Exporting and Importing Countries of a Non-Renewable Natural Resource: Taxation and Capturing Rents”, Emilio Cerdá y Xiral López-Otero.
References
- 2011-01: “Politicians' Luck of the Draw: Evidence from the Spanish Christmas Lottery”, Manuel F. Bagues y Berta Esteve-Volart.
References
- 2010-31: “The Effect of Family Background on Student Effort”, Pedro Landeras.
References
- 2010-29: “Random–Walk–Based Segregation Measures”, Coralio Ballester y Marc Vorsatz.
References
- 2010-28: “Incentives, resources and the organization of the school system”, Facundo Albornoz, Samuel Berlinski y Antonio Cabrales.
References
- 2010-27: “Retirement incentives, individual heterogeneity and labour transitions of employed and unemployed workers”, J. Ignacio García Pérez, Sergi Jimenez-Martín y Alfonso R. Sánchez-Martín.
References
- 2010-26: “Social Security and the job search behavior of workers approaching retirement”, J. Ignacio García Pérez y Alfonso R. Sánchez Martín.
References
- 2010-25: “A double sample selection model for unmet needs, formal care and informal caregiving hours of dependent people in Spain”, Sergi Jiménez-Martín y Cristina Vilaplana Prieto.
References
- 2010-24: “Health, disability and pathways into retirement in Spain”, Pilar García-Gómez, Sergi Jiménez-Martín y Judit Vall Castelló.
References
- 2010-23: Do we agree? Measuring the cohesiveness of preferences”, Jorge Alcalde-Unzu y Marc Vorsatz.
References
- 2010-22: “The Weight of the Crisis: Evidence From Newborns in Argentina”, Carlos Bozzoli y Climent Quintana-Domeneque.
References
- 2010-21: “Exclusive Content and the Next Generation Networks”·, Juan José Ganuza and María Fernanda Viecens.
References
- 2010-20: “The Determinants of Success in Primary Education in Spain”, Brindusa Anghel y Antonio Cabrales.
References
- 2010-19: “Explaining the fall of the skill wage premium in Spain”, Florentino Felgueroso, Manuel Hidalgo y Sergi Jiménez-Martín.
References
- 2010-18: “Some Students are Bigger than Others, Some Students’ Peers are Bigger than Other Students’ Peers”, Toni Mora y Joan Gil.
References
- 2010-17: “Electricity generation cost in isolated system: the complementarities of natural gas and renewables in the Canary Islands”, Gustavo A. Marrero y Francisco Javier Ramos-Real.