‹ Volver a la ficha Doc. eee-157

ESTUDIOS SOBRE LA ECONOMÍA ESPAÑOLA

EEE 157

February 2003

Figura

FEDEA Fundación de Estudios de Economía Aplicada

http://www.fedea.es/hojas/publicado.html

Carmen García Prieto Angel Martín Román Carlos Pérez Domínguez

February 2003

The returns to formal schooling in Spain are estimated in this paper. The main difference between this and previous papers on this subject is that, here, a distinction is made between the increase in the worker’s potential maximum wage due to schooling and the actual registered increase. This difference (or underpayment) can be justified on the basis of job search theory. We use the stochastic frontiers technique because it allows the estimation of variables (such as the potential wage) that cannot be directly observed.

One of the main results of this paper is that formal schooling clearly increases a worker’s potential maximum wage. This increase is particularly noticeable for those workers who have completed at least a five-year university programme. It has also been estimated that schooling increases the degree of underpayment, which is also quite relevant in the case of long-term university education. In spite of this, the effect of formal schooling on actual wages is clearly positive.

JEL Codes: I21, J24, J31. Key Words: human capital, labor income, stochastic frontiers.

Contact: Carmen García Prieto Dpto. Fundamentos del Análisis Económico. Facultad de Ciencias Económicas y Empresariales. Avda. Valle Esgueva 6. 47011 Valladolid. Spain e-mail: cgp@eco.uva.es tfno: +34 983 184428

The authors would like to thank the ‘Ministerio de Trabajo y Asuntos Sociales’ for their financial support.

1.- INTRODUCTION.

The analysis of the relationship between an individual’s level of formal schooling and their labor income has been of concern to economists for a long time. The work of Cantillon (1755) justifies the payment of higher wages to workers with better qualifications. Still within the classical period, Smith (1776) takes up these ideas once more. It could be said that he is the most direct predecessor of the modern theory of human capital developed at the beginning of the 1960s with the pioneering work of Schultz (1961), Mincer (1962) and Becker (1964).

The systematic analysis of the effect of schooling on labor income has a point of reference, which is fundamental to recent research on the subject: the work of Mincer (1974). The importance of this work is such that, thereafter, the most orthodox income equations have been called Mincer equations. The economic literature derived from this seminal work is ample. The estimation of the returns provided by schooling, based on the econometric adjustment of Mincer equations, is a topic that has given rise to much research over recent years as quality micro-databases have become more generally available in nonanglo-saxon countries .

The work of Griliches (1977) looked at some problems that could appear when estimating Mincer equations using ordinary least squared (OLS). One of the most frequent criticisms of the OLS estimation is that an individual’s schooling (which is one of the explanatory variables of labor income) is an endogenous variable. This would suggest the use of instrumental variables (IV) econometric methods.2 Willis & Rosen (1979) pose a different econometric problem: the very samples used to estimate earnings equation may not fulfill the basic requirement of being representative of the whole population. This is known in the literature on the subject as the self-selection bias.

Although this paper is closely related to the entire bibliography on this topic, in the sense that it takes the work of Mincer (1974) as its starting point, there are differences of both a methodological and conceptual nature which distinguish it from prior economic literature. To our understanding, this paper introduces an original element that sets it apart from other, previous papers. This original element lies in the estimation, not only of the effective returns obtained by individuals from their educational qualifications, but also of the potential returns associated with the various schooling levels. What is the reason for this difference between an individual’s potential wage and the actual wage earned? Because in the job search process workers do not have perfect information, and acquiring these information concerning employment opportunities has a cost. This is the theoretical framework of the so-called jobsearch theory associated with the work of McCall (1970), Mortensen (1970) and Lippman & McCall (1976a & 1976b). According to this theory, individuals fix a critical wage (the reservation wage) and when they get an offer of employment with an associate wage higher than the said critical value, they accept it. This supposes that many individuals will end their search before achieving the maximum wage they could aspire to, given their level of schooling. The difference between that ‘potential maximum’ wage and the effective wage earned is what is called ‘underpayment’.

1 Some recent works of research on this subject in the Spanish economy are: Alba-Ramírez & San Segundo (1995), De la Rica & Ugidos (1995), San Segundo (1997), Vila & Mora (1998), Barceinas et al. (2000), García et al. (2001) and Pons & Gonzalo (2001).
2 In the Spanish case, in accordance with the research of Barceinas et al. (2002), the returns of schooling obtained with IV estimations are very similar to those obtained from the OLS estimations when the sample is adequately screened.

The aim of this paper is, precisely, to measure that underpayment using Spanish data, and to see whether the difference between the potential and the effective wage increases in line with the individual’s level of schooling, or whether the opposite is true. The econometric technique of stochastic frontiers is used to achieve this aim. This technique has habitually been used in the framework of studies concerning productive efficiency, ever since Aigner et al. (1977) and Meeusen & Van den Broeck (1977), defined and used the concept of stochastic frontier in a simultaneous yet independent way.

Nevertheless, there are some applications of this technique in labor economics, to be more precise, in the wage setting process. As far as we know, the pioneering paper was that of Robinson & Wunnava (1989). The technique of stochastic frontiers is used in this paper to measure the female wage discrimination. Other prior papers were those of Hofler & Murphy (1992 and 1994). The first of them estimates up to what point workers achieve an effective wage below the potential maximum they could earn, given their marginal productivity (that is, the wage inefficiency is measured); while the second estimates the worker’s reservation wage. Both papers take the framework of the search theory as their setting. McClure, Girma & Hofler (1998) once more take up the question of wage inefficiency, but this time comparing stochastic frontier estimations for the United States and Canada. Polacheck & Robst (1998) take a similar line. Lang (2000) analyzes the question of wage discrimination among

German workers from different ethnic origins. Finally Watson (2000) relates wage inefficiency and the minimum wage in the United Kingdom, and deduces that the said minimum wage is not an effective economic policy in the fight against poverty.

The first new element of our work lies, precisely, in the use of this technique to measure the returns of schooling. It should also be pointed out that, despite the fact that there is already some literature on the subject of underpayment and labor income, we believe it to be the first time –and not only in Spain– that an analysis of how the latter is related to the different levels of schooling has been approached. Finally, it should be said that the use of this technique is not only a methodological novelty in the estimation of educational returns, but also that it has interesting properties from the strictly economic point of view. As shall be seen later, in addition to giving a measurement of the potential maximum wage an individual could reach with a particular level of schooling, it will also allow us to measure the efficiency of the employment search process of each educational group.

The rest of the paper is organized as follows: Section two offers the theoretical basis on which our estimations are supported. Section three gives the econometric specifications of the model, pointing out, firstly, the estimation technique used, and secondly, the data used. Section four gives the results obtained in the estimation, while section five, the last, summarizes the main conclusions. The paper ends with four appendices: the first is of a technical nature; the second explains the variables used and shows some statistics describing the sample used; the third shows the complete results of the estimations; and the fourth shows the validation test.

2.- ACTUAL AND POTENTIAL EARNINGS: THEORETICAL BASIS

As a starting point, let us suppose that there is an individual function for generating potential income as described by the following equation:

\[w _ {i} ^ {P} = f (X _ {i})\tag{1}\]

It is a technological relation that determines the maximum wage earnings that the ith worker can obtain given certain income generating inputs represented by the vector The said inputs are basically determined by the human capital the worker possesses3. It should be pointed out that expression (1) presupposes an ‘efficient’ behavior of the worker, in the sense that it fixes the maximum amount of money the worker could earn with the best use of his/her formal knowledge and work tenure. In other words, the above expression is an upper boundary for labor income, so the actual or effective income earned by the worker at any given time ) must be lower than, or at best equal to, the maximum potential

The main cause of underpayment, that is, that workers may not actually be earning their potential wage, is to be found in the existence of imperfect information. Thus, the job search process becomes costly for the worker. In this way, for a particular individual looking for a job, the best option may be to accept a post offering a wage below his/her maximum potential. This will always be so as long as the marginal cost of continuing to search for employment exceeds the expected marginal benefit of searching.4

To be more precise, the reasons why a certain individual’s effective wage falls below the maximum potential, or wage frontier, can be put into two main categories.

The first of these categories has an essentially ‘objective’ nature. It is related to the wage distribution that a particular individual has to face when looking for a job, given his/her earnings generating inputs. It would be expected that workers with the lowest level of formation and tenure would have to face a concentrated wage distribution; that is, that their possible wage range will be fairly restricted. However, as the worker’s human capital increases, so will her/his wage possibilities. Thus, individuals with the least qualifications, who face a concentrated wage distribution, should, on average, find themselves closer to their potential wage. A good example is the situation of a teenager, who can only realistically expect to earn the legal minimum wage. It is thus very likely that the wage frontier for such a worker will be very close to the minimum wage and that the effective wage of most workers in this collective is very close to the said frontier.

3 In accordance with the proposals of Mincer (1974)
In accordance with the Job Search Theory, the worker’s optimum strategy is to determine a reservation wage in such a way that any job offer with a wage below it is rejected, while the first offer of employment providing a wage equal to or higher than the reservation wage is accepted. In order to determine this acceptance wage, the worker must take into account precisely those costs and benefits associated with fixing a marginally higher reservation wage. See McCall (1970), Mortensen (1970) and Lippman & McCall (1976a & 1976b).

Other reasons that influence the wage distribution an individual must face concern the structure of the local labor market, that is, its professional and industrial structures.

The second of these categories is related to more ‘subjective’ questions. On the one hand, there are all those factors that determine an individual’s reservation wage, and on the other, the elements associated with their efficiency in searching for employment. Given the same wage distribution, those with higher acceptance wages and who carry out the search mechanisms with greater efficiency will earn wages closer to their maximum potentials. For instance, let us consider two individuals A and B with identical ‘objective characteristics”. Let us suppose that A is married, that the spouse works and that they have no children, while B is divorced or married, but that the spouse is unemployed, and that they have children. It would be expected that the individual A would fix a higher reservation wage (he or she will be more demanding in accepting employment), since there is a subsidiary income (that of the spouse) and because they have no children. On the other hand, the individual B would fix a much lower reservation wage for the reasons inversely opposite to those of A. The conclusion is that A, with a high probability, will be closer to their wage frontier than B, simply because he/she has been more demanding when selecting wage offers.

Figure 1 shows graphically the arguments expressed above. The abscissa measures wages and the ordinate measures the number of posts the individual can find for each pre-fixed wage level (thus, what it represents is not exactly a function of density). Two “bells” can be seen. The smaller one (solid line) is associated with the situation of an individual with few income-generating inputs (an unqualified youth). The larger “bell” (dotted line) corresponds to a worker with greater wage possibilities (an older person with a higher level of schooling). Such a figure shows the fact that an older person with better qualifications can always carry out the work of an unqualified youth, while the opposite is not true. The potential wage of the person with a higher level of schooling, , is greater than that of the youth, , and it is this that the estimation of our wage frontier shows.

Let us imagine that the ‘subjective’ circumstances of both individuals are such that they fix their reservation wages around the mean for the corresponding distributions for the youth and for the qualified adult). The above supposition incorporates the fact that the person with a higher level of schooling, aware of the fact that she/he has access to a greater variety of job offers, will fix a reservation wage which is greater than that of the young uneducated person. However, the distance to the corresponding maximum potential is smaller (even in relative terms) in the case of the youth than in the case of the person with a higher level of schooling.

3.- ECONOMETRIC TECHNIQUE

Estimation method

The stochastic frontier estimation techniques offer a plausible means for the estimation of the potential wage a worker could earn.

The method basically consists in considering that the effective wage of each individual is equal to, or lower than, the maximum level that can be reached on the market (the potential wage). The potential wage forms the upper limit of the observations and is obtained, as we have already seen, from a set of variables that give an estimate of the marginal productivity of each individual.

Let be the effective wage earned by the ith worker. We suppose that this can be explained by the following model:

\[\log w _ {i} = \log w _ {i} ^ {p} - u _ {i}\tag{2}\]

where is the potential wage and a random, non-negative disturbance.

The potential wage of each individual, which is their frontier, is obtained from a set of variables all of them reflecting their income generating inputs (basically, their human capital), according to the following specification:

\[\log w _ {i} ^ {p} = \beta^ {\prime} X _ {i} + v _ {i}\tag{3}\]

where is a parameter vector to be estimated, is the vector of the income generating inputs, and a random disturbance term that gives the frontier a stochastic nature.

Substituting (3) in (2), we get:

\[\log w _ {i} = \beta^ {\prime} X _ {i} + v _ {i} - u _ {i}\tag{4}\]

We explain the difference between the potential wage and the effective wage a worker earns by using a set of variables specific to each individual5, and a new random disturbance term , in accordance with:

\[u _ {i} = \delta^ {\prime} z _ {i} + \xi_ {i}\tag{5}\]

where is a parameter vector to be estimated.

To guarantee that , we consider that is distributed identically among the sample as a normal with zero mean and variance , truncated at the point , in such a way that, . Thus, is distributed as a normal variable whose mean depends on the specific explanatory variables of the individuals and truncated at zero, . On the other hand, we suppose that is distributed , independent of and of the regressors. Thus, we have a model with a composed error , whose likelihood function, considering the existence of N individuals in the sample, is shown in Appendix 1.

Using maximum likelihood, we get consistent estimators of the frontier parameters, so individuals’ potential wage estimation is consistent. The estimators of the parameters that accompany the explanatory variables of the underpayment are also consistent. However, the estimation does not give a value for every , as this is integrated in the compound error term . To find a specific value for each individual of the amount that separates her/him from their potential wage, it is necessary to consider the conditioned density function

5 Following the specification of Huang & Liu (1994) and Battese & Coelli (1995).
6 In some papers –e.g. Lang (2000)– the previous estimations are carried out in two stages. The determining parameters of the frontier are estimated in the first, while those of the inefficiency are estimated in the second. This procedure is inconsistent (see, for instance Kumbhakar and Lovell, 2000).

Data used and description of the variables

The data source used to carry out the estimations was the Household Panel of the European Union (PHOGUE) in its third edition, corresponding to the year 1996, with data referring to Spain, which offers individualized information on 15,643 people and 6,268 households. We selected from the sample, wageearning males working 15 or more hours per week , thus reducing the number of observations to a total of 2,780. Likewise, the data concerning the households of these individuals was also used to obtain some variables.

The sample of wage-earning males used gives us a homogeneous and numerous group in the population which minimizes possible problems of selfselection bias. As suggested by San Segundo (1997), bias can come from the use of a sample made up of workers from both sexes9 or because of the incorporation of self-employed workers whose reported income can be less reliable than wage-earners. Workers clocking up few hours per week have also been excluded from the chosen sample.

7 This was proposed by Jondrow et al. (1982). It would be possible to use the mean or the mode of this conditioned distribution to obtain an unbiased, though inconsistent estimator of the underpayment.
8 The PHOGUE offers two ways of finding out the hours worked by an individual: their personal statement and the objective classification technique suggested by the Current Population Survey (Spanish National Institute of Statistics). We have chosen the second option.
9 Pons & Gonzalo (2001), also using Spanish data, justify the use of an exclusively male sample in order to avoid the possible complications derived from the fact that, with females, many professional careers are interrupted by the birth and care of children.

The dependent variable considered is the log of the net hourly wage.10 The variables of human capital used for the estimation of the wage frontier (Xi) refer to the level of formal schooling reached by the worker, his/her labor tenure and age.11 The variables used to pick up the difference between the real and the potential wage of each worker (zi) refer to their age and level of schooling, their marital status and dependent relatives, the individual’s non-labor income, the geographical area of residence (NUT), the industry in which the individual works and their labor mobility capacity.

Appendix 2 gives a detailed explanation of the way in which all these variables were made, as well as a summary of the descriptive statistics corresponding to the variables associated with schooling.

4.- RESULTS

Table 1 shows the results obtained in the different estimations we have carried out concerning the variables of formal schooling. Appendix 3 shows the complete set of results. The estimation of the stochastic frontiers was carried out using the computer program FRONTIER 4.1, developed in the Centre for Efficiency and Productivity Analysis (CEPA) of the University of New England (Australia) to study the productive efficiency through frontier functions. A brief guide to how it works can be found in Coelli (1996).

[Insert Table 1]

Appendix 4 shows the results of the validation tests of the estimated models, the main results of which are summarized below.

Columns (I) and (II) of Table 1 shows the results of the estimations made using Ordinary Least Squared which constitute our first reference point. The first of these columns incorporates the number of years of schooling as a continuous variable, obtaining a mean annual return to schooling of 5%. On the other hand, the estimation of column (II) incorporates schooling in a discrete way; in this case, the coefficients show the growth rate of the wage associated to each stage of formal schooling, taking the least qualified group (less primary) as the reference point. The growth of the rate as the subject covers higher levels of schooling can be appreciated. In both cases the estimated returns are in line with those obtained in prior research carried out for Spain.12

10 The said variable is made up of the quotient between the current net monthly income derived from work as an employee and the mean number of hours worked per month. The use of net wages in our study must be stressed. This fact could explain the slightly lower value of the estimated returns with respect to other papers which use the gross wage, as is the case with Barceinas et al. (2002).
11 These last two variables try to reflect the worker’s specific and generic experience.

Columns (III) to (VIII) show the results of the estimations of the stochastic frontiers. In the case of column (III) it can be seen how the maximum potential wage increase associated with an extra year’s formal schooling rises to 5.6%. However, the underpayment (difference between the real and potential wage) also increases with schooling by 1% per year –column (V)-, that is, each year of formal schooling separates the maximum potential wage from the effective wage earned by that percentage. As a result of both phenomena, the effective mean wage of a worker grows annually by 4.6% -column (VII)– a figure slightly below that obtained in the OLS estimations.

The results obtained on incorporating levels of schooling as a discrete variable are graphically summarized in figure 2. The continuous growth of the potential returns corresponding to higher levels of schooling and, most especially, that associated with ‘long cycle’ university studies can be appreciated –column (IV) of table 1–. However, the degree of underpayment – column (VI)– shows a less regular behavior although, in general, it can be seen to be reduced in the groups with the lowest levels of schooling and noticeably higher in the case of ‘long cycle’ university studies. As a result of these two effects, the increases in the effective mean wage is shown in column (VIII); as it can be seen, in general, the values are slightly lower than those obtained from the OLS estimation –column (II)–.

[Insert Figure 2]

To conclude with the analysis of the results we have put together Table 2, based on the estimations of the stochastic frontier with discrete levels of schooling. It shows information on the estimated hourly wage levels for each schooling group. Column IV of the said table calculates the percentage of wage achievement for each group as the quotient between the ‘de facto’ earned wage –column II– and the potential maximum –column I–. It is interesting to see that the said percentage reaches its lowest value for the collective with the highest level of schooling. An individual belonging to this group achieves, on average, 73% of their maximum potential wage, while, for the rest of the groups, the percentage is, on average, equal to or higher than 80%. The dispersion of this wage achievement is also significantly higher in the group with the highest level of schooling, as shown by the variance, which reveals the presence in this collective of a wide range of professions with varying remuneration.

12 See, for instance, the work of San Segundo (1997) and Pons & Gonzalo (2001).

[Insert Table 2]

Looking further into this phenomenon, it can be appreciated how the worst paid individual with ‘long cycle’ university studies hardly reaches 12% of their potential wage, a value much lower than that corresponding to the rest of the schooling groups. Nevertheless, the degree of wage achievement of the bestpaid individual hardly differs in each respective group. This result may be related to the phenomenon of over-education, which is especially important in Spain13.

Finally, it should be pointed out that the lower percentage of wage achievement by the highest qualified individuals is also true for all the deciles of the distribution, the difference being more acute the lower the decile being considered.

5.- SUMMARY AND CONCLUSIONS

The returns to formal schooling in Spain are estimated in this paper. The main difference between this and previous papers on this subject is that, here, a distinction is made between the increase in the worker’s potential maximum wage due to schooling and the actual registered increase. This difference (or underpayment) can be justified on the basis of job search theory

The stochastic frontiers technique was used to carry out the estimations because it allows the approximation of the potential wage, a variable that cannot be directly observed.

The main results obtained are as follows. The increase in the potential wage associated with an extra year of formal schooling is 5.6%. However, the underpayment (difference between the real and the maximum potential) also increases with schooling by 1% per year, that is, each year of formal schooling separates the maximum potential wage from that ‘de facto’ earned by this percentage. As a result of these phenomena the worker’s real mean wage increases annually by 4.6%.

13 On this subject, see Dolado et al. (2002) and the related bibliography mentioned there.

If we separate schooling into discreet sections (in accordance with the highest qualification attained by the worker) the increase in the returns associated with higher levels of schooling and, most especially, that corresponding to ‘long cycle’ university studies can clearly be seen. The degree of underpayment shows, in this case, a less regular behavior. Nevertheless, it can be seen that the magnitude is smaller in the groups with the lowest qualification and noticeably higher in the case of ‘long cycle’ university studies. In spite of this, the effect on the effective wage of a higher level of schooling is clearly positive.

TABLE 1: Schooling Equation

OLSPotential WageUnderpaymentActual Wage
(I)(II)(III)(IV)(V)(VI)(VII)(VIII)
Years of schooling5.0%5.6%1.%4.6%
Primary4.0%10.0%6.9%2.9%
Secondary (1st. Level)16.6%23.7%8.9%13.5%
FP (1st. Level)19.0%35.2%16.8%15.7%
FP ( $2^{nd}$ Level)33.8%41.8%11.0%27.8%
Secondary (2nd Level)39.5%49.9%12.3%33.5%
University (Short cycle)82.%82.9%5.3%73.7%
University (Long cycle)96.3%145.5%26.5%94.1%

• Source: Appendix 3.

The coefficients associated with the different schooling levels represent growth rates respect to the reference group “less primary and no schooling”. These rates have been elaborated in accordance with the following expression: exp(β)-1, where β is the coefficient obtained in the estimation (see appendix 3).

The values of the underpayment and the actual wage are calculated from the estimations of appendix 3 in accordance with the specifications of appendix 1.

TABLE 2: Underpayment by schooling

Wage / hourDegree of wage achievement (II) / (I)
Mean Potential Wage (I)Mean Actual Wage (II)=(I)-(III)Mean Underpay. (III)MeanVar.Min.Max.10%20%30%40%50%60%70%80%90%
Less primary839.56708.55131.010.850.0080.500.950.710.810.830.860.880.890.900.910.92
Primary897.16735.65161.510.820.0120.250.960.690.770.810.840.850.870.890.900.92
Secondary (1st. Level)901.20747.79153.410.830.0090.310.950.710.780.810.840.860.870.890.900.92
FP ( $1^{st}$ Level)976.90773.93202.970.790.0170.240.940.660.720.750.800.820.850.870.890.91
FP ( $2^{nd}$ Level)1023.14856.80166.340.840.0080.260.960.730.770.820.840.860.870.890.900.92
Secondary (2nd Level)1175.84975.25200.580.830.0090.230.950.700.760.800.830.850.870.880.900.92
University (Short cycle)1582.391371.13211.270.870.0050.460.950.770.830.860.880.890.900.910.920.93
University (Long cycle)2046.521491.32555.200.730.0250.120.940.510.610.680.740.770.800.830.860.88

• Source : Frontier estimations • (I) Mean estimated potential wage for each group of schooling • (III) Mean estimated underpayment for each group of schooling

Figura

Returns to schooling by educational level

Returns to schooling by educational level

The likelihood function used to obtain the frontier estimation is as follows:

\[\ln L = c t e - \frac {N}{2} \ln \sigma^ {2} - \sum_ {i = 1} ^ {N} \ln \Phi \left(d _ {i}\right) + \sum_ {i = 1} ^ {N} \ln \Phi \left(d _ {i} ^ {*}\right) - \frac {1}{2} \sum_ {i = 1} ^ {N} \frac {\left(\varepsilon_ {i} + \delta^ {\prime} z _ {i}\right) ^ {2}}{\sigma^ {2}}\]

' iz where ( )2id = and , been , and

is approaching the total residual variance of , and indicates the relative contribution of to this residual variance.

Method to obtain the elasticities:

Following Huang and Liu (1994), the effect on the individual expected effect of a variable that simultaneously explains both the frontier and the

underpayment

\[\frac {\partial \ln w _ {i}}{\partial x _ {i j}} = \beta_ {j} - \psi_ {i} \frac {\partial (\delta^ {\prime} Z _ {i})}{\partial x _ {i j}},\]

and Φ are the standard normal

density and cumulative distribution functions. This expression is composed of two terms: the effect of the variable on the frontier, , and the effect on the underpayment,

APPENDIX 2: Description of the variables used in the estimation

Sex: Dummy variable that takes the value one if the worker is a woman and zero if it is a man.

Age: Worker’s age.

AgeSq: Square of the worker’s age.

Ten: Tenure, number of years the worker has been employed in his/her present job.

TenSq: Tenure squared.

Ed1: Dummy variable that takes the value one if the worker is illiterate or has no studies (2 years) and zero otherwise.

Ed2: Dummy variable that takes the value one if the worker completed primary school (5 years) and zero otherwise.

Ed3: Dummy variable that takes the value one if the worker has completed a first level of secondary education (8 years) and zero otherwise.

Ed4: Dummy variable that takes the value one if the worker has completed a cycle of further education (9 years) and zero otherwise.

Ed5: Dummy variable that takes the value one if the worker has completed a second cycle of further education (11 years) and zero otherwise.

Ed6: Dummy variable that takes the value one if the worker has completed a second level of secondary education (12 years) and zero otherwise.

Ed7: Dummy variable that takes the value one if the worker has a shortcycle university degree or equivalent (15 years) and zero otherwise.

Ed8: Dummy variable that takes the value one if the worker has a longcycle university degree or equivalent (17 years) and zero otherwise.

Marrdep: Dummy variable that takes the value one if the worker is married and has, at least, a dependent child, and zero otherwise.

Marrnodep: Dummy variable that takes the value one if the worker is married and has not dependent children, and zero otherwise.

Singledep: Dummy variable that takes the value one if the worker is single and has, at least, a dependent child, and zero otherwise.

Famincome: Total income of a household where the worker lives, minus that worker’s labour income.

Nutma: Dummy variable that takes the value one if the worker lives in Madrid and zero otherwise.

Nutca: Dummy variable that takes the value one if the worker lives in Canarias and zero otherwise.

Nutce: Dummy variable that takes the value one if the worker lives in Castilla y León, Castilla-La Mancha or Extremadura and zero otherwise.

Nutes: Dummy variable that takes the value one if the worker lives in

Cataluña, Comunidad Valenciana or Baleares and zero otherwise.

Nutne: Dummy variable that takes the value one if the worker lives in País Vasco, Navarra, La Rioja or Aragón and zero otherwise.

Nutno: Dummy variable that takes the value one if the worker lives in Galicia, Asturias or Cantabria and zero otherwise.

Nutsu: Dummy variable that takes the value one if the worker lives in Andalucía, Murcia, Ceuta or Melilla and zero otherwise.

Agricult: Dummy variable that takes the value one if the worker’s professional activity is in the farming sector and zero otherwise.

Energy: Dummy variable that takes the value one if the worker’s professional activity is in the energy and mining industries and zero otherwise.

Manufact: Dummy variable that takes the value one if the worker’s professional activity is in the manufacturing sector and zero otherwise.

Mineral: Dummy variable that takes the value one if the worker’s professional activity is in the metallic and non-metallic mineral sector and zero otherwise.

Machinery: Dummy variable that takes the value one if the worker’s professional activity is in the machinery and equipment sector and zero otherwise.

Construction: Dummy variable that takes the value one if the worker’s professional activity is in the construction sector and zero otherwise.

Saleserv: Dummy variable that takes the value one if the worker’s professional activity is in the sector of sales services and zero otherwise.

Finaserv: Dummy variable that takes the value one if the worker’s professional activity is in the sector of banking and financial services and zero otherwise.

AAPP: Dummy variable that takes the value one if the worker’s professional activity is in the Public Administration and Social Services sector and zero otherwise.

Eduhealth: Dummy variable that takes the value one if the worker’s professional activity is in the education and health care sector and zero otherwise.

Others: Dummy variable that takes the value one if the worker’s professional activity is in other sectors of the NACE and zero otherwise.

Immobility: Dummy variable that takes the value one if the worker was born in Spain and has resided in the same region ever since and zero otherwise.

Main descriptive statistics

MeanMax.Min.St. DeviationN obs.
Wage931,355576,3281,72516,212780
Age38,7669,0017,0011,152780
Tenure7,8315,000,006,202780
Education8,9717,002,004,122780
Wage by educational level
Less primary725,561953,53250,33266,27131
Primary759,543037,3884,42315,19794
Secondary (1st. Level)780,502955,61116,82358,34647
FP ( $1^{st}$ Level)798,002076,84108,22336,88214
FP ( $2^{nd}$ Level)906,923119,16245,33446,27226
Secondary (2nd Level)1034,693178,41132,86504,53319
University (Short cycle)1431,202951,30400,53534,04179
University (Long cycle)1570,575576,3281,72777,92270

Ordinary Least Squares (OLS)

Coefficient St. Deviation. t-statisticCoefficient St. Deviation. t-statistic
Intercept4,9940,08657,8175,2650,09257,031
Age0,0450,0049,9420,0440,0049,678
Age Sq-0,0000,000-7,8805-0,0005,481-8,012
Tenure0,0390,0057,2670,0410,0057,515
Tenure Sq-0,0000,000-2,782-0,0000,000-2,836
Education0,0500,00128,793
Ed20,0390,0351,130
Ed30,1530,0364,175
Ed40,1730,0424,092
Ed50,2910,0416,933
Ed60,3330,0398,407
Ed70,5990,04313,922
Ed80,6740,04016,865
$R^2$ 0,4510,457
DW1,9761,990
$σ^2$ 0,3680,366
Log-likelihood-1168,832-1150,921

Stochastic Frontiers

Coefficient St. Deviation. t-statisticCoefficient St. Deviation. t-statistic
FRONTIER
Intercept5,3430,08761,0865,5810,09558,861
Age0,0350,0047,8540,0340,0047,541
Age Sq0,0000,000-5,1910,0000,000-5,214
Tenure0,0340,0056,5670,0340,0056,681
Tenure Sq-0,0010,000-2,262-0,0010,000-2,056
Education0,0560,00226,577
Ed20,0950,0392,414
Ed30,2120,0415,134
Ed40,3010,0496,125
Ed50,3490,0506,984
Ed60,4050,0458,914
Ed70,6040,05211,700
Ed80,8980,05416,744
UNDERPAYMENT
Intercept-2,4820,627-3,957-2,5840,758-3,411
Age0,0290,0064,4770,0290,0074,345
Age Sq-0,5120,107-4,778-0,4690,093-5,050
Marrnodep-0,4300,118-3,647-0,3930,095-4,137
Singledep0,2140,0772,7830,2130,1101,943
Famincome0,0000,000-0,3200,0000,0000,132
Education0,0780,0135,825
Ed20,4910,2292,142
Ed30,6410,2522,540
Ed41,1630,3393,431
Ed50,7820,3202,446
Ed60,8650,2803,095
Ed70,3850,4030,955
Ed81,7690,3894,553
Nutca0,5010,1273,9540,4980,1194,186
Nutce0,3800,1093,4990,3460,0973,581
Nutes0,1620,0792,0480,1110,0911,216
Nutne-0,0580,110-0,529-0,0710,081-0,880
Nutno0,7320,1524,8240,7130,1544,640
Nutsu0,4670,1223,8240,4750,1213,927
Agricult0,5900,1175,0390,6670,1594,204
Energy-3,0631,422-2,154-2,7181,802-1,508
Manufact-0,1540,078-1,978-0,1510,070-2,171
Mineral-0,6350,149-4,247-0,6670,151-4,406
Machinery-0,7590,215-3,538-0,8040,223-3,604
Construction-0,6700,178-3,767-0,6520,208-3,129
Finaserv-0,8490,203-4,181-0,8450,219-3,866
AAPP-1,0660,250-4,260-0,9800,254-3,861
Eduhealth-1,2710,329-3,858-0,9160,217-4,230
Others-0,3480,098-3,555-0,3460,111-3,126
Immobility0,2580,0673,8370,2470,0643,882
$\sigma^2$ 0,3600,0645,5790,3490,0585,996
$\gamma$ 0,7720,04019,0560,7690,03721,061
Log-likelihood-980,107-946,838
LR-test of one-side error377,450408,164

The following hypothesis are established:

• We have considerer a simpler model in wich the underpayment is not dependent on the variables suggested. In such a way, the ui component would have a constant mean, equal for all the individuals. The hypothesis will be when education is a continuous variable, and when it is discrete.

• If the mean is equal to cero, the hypothesis will be when education is a continuous variable, and when it is discrete.

Finally, we contrast the existence of a frontier, . In this way, the variables that explain the underpayment will became explicative of the effective wage, and the estimation becomes an OLS regression. The hypothesis will be in this case H0: when education is a continuous variable, and when it is discrete, because there is an intercept in the frontier as well as age and education variables.

All these hypothesis are rejected as it can be seen from the next table:

14 Where, indicates the relative contribution of u to the total residual i variance.

Education as a continuos variable:

NULL HYPÓTHESISLOGLIKELIHOODSTATISTIC (*)DECISIÓN
$H_0: \delta_1=\delta_2=...=\delta_{23}=0$ -1120.6281.1Reject
$H_0: \delta_0=\delta_1=...=\delta_{23}=0$ -1137.8315.5Reject
$H_0: \gamma=\delta_0=\delta_1=\delta_6=0$ -1026.592.8Reject

Education with dummies

NULL HYPOTHESISLOGLIKELIHOODSTATISTIC (*)DECISIÓN
$H_0: \delta_1=\delta_2=...=\delta_{29}=0$ -1101.5309.2Reject
$H_0: \delta_0=\delta_1=...=\delta_{29}=0$ -1119.5345.3Reject
$H_0: \gamma=\delta_0=\delta_1=\delta_6=...=\delta_{12}=0$ -1006.6119.4Reject

(*)The statistic was calculated for all cases, as λ=-2[loglikelihood(H0)-loglikelihood . It is distributed as a with as many degrees of freedom as parameters considered to be zero in the null hypothesis. In the last test, the statistic is distributed as a mix of , and the critical values can be found in Kodde and Palm (1986). If it works out to be greater than the value given by the tables at 95%, then the hypothesis is rejected.

References

  1. Aigner, D.J., Lovell, C.A.K. and Schmidt, P.J. (1977) Formulation and Estimation of Stochastic Frontier Production Function Models, Journal of Econometrics, 6, pp. 21-37.
  2. Alba, A. and San Segundo, M.J., (1995) The Returns of Schooling in Spain, Economics of Schooling Review, 14, pp. 155-166.
  3. Barceinas, F. Oliver, J., Raymon, J.L. and Roig, J.L. (2000) Spain. In: Harmon, C., Walker, I. and Westergrad-Nielsen, N. (Ed.), Schooling and Earnings in Europe: a Cross Country Analysis of the Return to Schooling (E. Elgar).
  4. Barceinas, F. Oliver, J., Raymon, J.L. and Roig, J.L. (2002) Los rendimientos de la educación en España, mimeo. http://www.minhac.es/ief/Seminarios/EconomiaPublica/20020214.pdf
  5. Battese, G.E. and Coelli T. (1995) A Model for Technical Inefficiency Effects in a Stochastic Frontier Production Function for Panel Data, Empirical Economics, 20, pp. 325-332.
  6. Becker, G. S. (1964) Human Capital: a Theoretical and Empirical Analysis, with Special Reference to Schooling, New York: Columbia University Press.
  7. Cantillon, R. (1755) Essai sur la nature du commerce en général, translated by Henry Higgs, New York, Augustus Kelley, 1964.
  8. Coelli, T. (1996) A Guide to FRONTIER Version 4.1: A Computer Program for Stochastic Frontier Production and Cost Function Estimation, CEPA working paper 96/07.
  9. Dolado,J. J., Jansen, M. and Jimeno, J. F. (2000) A Matching Model of Crowding-Out and On-the-Job Search (with an application to Spain)”, Documento de trabajo FEDEA 02-16. FEDEA, Madrid.
  10. De la Rica, S. and Ugidos, A. (1995) ¿ Son las diferencias en capital humano determinantes de las diferencias salariales observadas entre hombres y mujeres?, Investigaciones Económicas, 19 (3), pp. 395-414.
  11. García, J., Hernández, P.J. and López A. (2001) How Wide is the Gap? An Investigation of Gender Wage Differences Using Quantile Regression, Empirical Economics, 26(1), pp. 149-167.
  12. Griliches, Z. (1977) Estimating the Return to Schooling: Some Econometric Problems, Econometrica, 45, pp. 1-22.
  13. Hofler, R. A and Murphy K., J. (1992) Underpaid and Overworked: Measuring the Effect of Imperfect Information on Wages, Economic Inquiry, 30 (3), pp. 511-529.
  14. Hofler, R. A and Murphy K., J. (1994) Estimating Reservation Wages of Employed Workers Using a Stochastic Frontier, Southern Economic Journal, 60, 4, pp.961-976.
  15. Huang, C. J. and Liu, J. T. (1994) Estimation of a Non-Neutral Stochastic Frontier Production Function, The Journal of Productivity Analysis, 5, pp. 171-180.
  16. Jondrow, J., Lovell, C.A.K., Materov, I.S. and Schmidt P. (1982) On the Estimation of Technical Inefficiency in the Stochastic Frontier Production Function Model, Journal of Econometrics, 19, 233-238.
  17. Kodde, D.A. y F.C. Palm (1986) Wald criteria for jointly testing equality and inequality restrictions, Econometrica, 54, 1243-1248.
  18. Kumbhakar, S. C. and Lovell, C. A. K. (2000) Stochastic Frontier Analysis, Cambridge University Press. New York
  19. Lang, G. (2000) Native-immigrant Wage Differentials in Germany - Assimilation, Discrimination, or Human Capital?, Universitaet Augsburg, Institute for Economics, discussion paper, n. 197
  20. Lippman, S. A. and McCall, J. J. (1976a) The Economics of Job Search: a Survey”, Part I, Economic Inquiry, 14, pp. 155-189.
  21. Lippman, S. A. and McCall, J. J. (1976b) The Economics of Job Search: a Survey”, Part II, Economic Inquiry, 14, pp. 347-368.
  22. McCall, J. J. (1970) Economics of Information and Job Search, Quarterly Journal of Economics, 84, 113-126.
  23. McClure, K. G., Girma, P. B. y Hofler, R. A. (1998) International Labour Underpayment: a Stochastic Frontier Comparison of Canada and the United States, Canadian Journal of Regional-Science,21(1), pp. 87- 109.
  24. Meeusen, W. and Van den Broeck, J. (1977) Efficiency Estimation from Cobb-Douglas Production Functions with Composed Error, International Economic Review 18, 2, pp. 435-444.
  25. Mincer, J. (1962) On-the-job Training: Costs, Returns and Some Implications, Journal of Political Economy, 70 (5), part 2.
  26. Mincer, J. (1974) Schooling, Experience, and Earnings, Columbia University Press. NBER, New York.
  27. Mortensen, D. T. (1970) Job Search, the Duration of Unemployment and the Phillips Curve, American Economic Review, 60, pp. 847-862.
  28. Polacheck, S.W. and Robst J. (1998) Employee Labor Market Information: Comparing Direct World of Work Measures of Workers’ Knowledge to Stochastic Frontier Estimates, Labour Economics, 5, 231-242.
  29. Pons, E and Gonzalo, M. T (2001) Return to Schooling in Spain. How Reliable Are IV Estimates?, IV Jornadas de Economía Laboral, Valencia.
  30. Robinson, M.D. and Wunnava, P. V. (1989) Measuring Direct Discrimination in Labor Markets Using a Frontier Approach: Evidence from CPS Female Earnings Data, Southern Economic Journal, 56, pp. 212-218.
  31. San Segundo, M J. (1997) Educación e ingresos en el mercado de trabajo español, Cuadernos Económicos de ICE, 63, pp. 105-123.
  32. Schultz, T. W. (1961) Invesment in Human Capital, American Economic Review, 51: pp. 1-17.
  33. Smith, A. (1776) An Inquiry into the Nature and Causes of the Wealth of Nations, eds. R.H Campbell and A.S Skinner (1976).
  34. Vila, L. and Mora, J.G. (1998) Changing Returns to Schooling in Spain During the 80’s, Economic of Schooling Review, 17 (2), pp. 173-178.
  35. Watson, D. (2000) UK Underpayment: Implications for the Minimun Wage, Applied Economics, 32, pp. 429-440.
  36. Willis, R. J. and Rosen, S. (1979) Schooling and self-selection, Journal of Political Economy, 87 (5), pp. 817-836.