‹ Volver a la ficha Doc. dt2021-15

Documento de Trabajo - 2021/15

Ángel de la Fuente (FEDEA and Instituto de Análisis Económico (CSIC)) Rafael Doménech (BBVA Research and Universidad de Valencia)

Noviembre 2021

fedea

Las opiniones recogidas en este documento son las de sus autores y no coinciden necesariamente con las de FEDEA.

Ángel de la Fuente FEDEA and Instituto de Análisis Económico (CSIC)

and Rafael Doménech BBVA Research and Universidad de Valencia

November 2021

Abstract

Scores in standardized international student achievement tests and some recent adult literacy studies provide interesting data on the quality of educational outputs and on the skill level of the population that can be a useful complement to the data on the quantity of schooling which have been most commonly used in the growth literature. This paper describes the most recent available primary data on the subject, reviews different attempts to organize, standardize and summarize them, and discusses the strengths and weaknesses of the existing indicators and their potential usefulness as explanatory variables in empirical analyses of the determinants of income and welfare levels and growth rates. A final section investigates the distribution of these indicators across a sample of 21 core OECD countries.

Keywords: human capital measurement, years of schooling, educational quality; adult skills JEL Classification: O4, I2

* The authors gratefully acknowledge financial support from the Spanish and Valencian governments through research projects CICYT PID 2020-116242RB-I00 and PROMETEO 2016-09.

1. Introduction

Most of the empirical literature on human capital and growth has relied on data on the quantity of education, often measured by the average number of years of schooling of the population. It is clear, however, that years of schooling can be at best an imperfect proxy for the stock of human capital, a limitation that can bias the estimation of its effects on economic and social progress (see, for example, Wößmann, 2003, Folloni and Vittadini, 2010, or Breton, 2011). Knowledge and skill levels will vary across countries with similar school attainments if there are differences among them in the quality of their educational systems or in the extent to which skills are built up or maintained through other channels, such as various types of post-school training and on-the-job learning. In recent years, researchers have become more keenly aware of the limitations of quantity of education variables and have paid increasing attention to the quality of education and to direct indicators of the skills and competences of the population.

This paper reviews the available cross-country data on skill levels and educational quality and analyzes their distribution across OECD countries and their strengths and limitations in comparison to years of schooling. Section 2 describes the available primary data from standardized international assessments of student and adult competences. Section 3 deals with different attempts to organize, standardize and summarize these data. Section 4 discusses the strengths and weaknesses of the existing indicators and their potential usefulness as explanatory variables in empirical analyses of the determinants of income and welfare levels and growth rates. Section 5 investigate the distribution of these indicators across a sample of 21 core OECD countries for which the quality of the data and the number of observations available are greater than for developing countries. Finally, section 6 presents the main conclusion of this selective survey.

2. Primary data on student achievement and on skill levels from standardized international tests

To approximate the quality of education in a country, researchers have generally relied on its performance in standardized international tests that measure the knowledge or competences of the student population, although there have also been some studies that have used estimates of Mincerian returns to schooling as quality indicators. Student achievement tests come in two varieties. The first one measures the academic achievement of students at different stages of their primary and secondary education, focusing on their mastery of standard curricula. The second set of tests is also administered to students in mandatory education but focuses on the command of the basic and applied skills that can be identified with a broad concept of literacy (and numeracy), rather than on academic achievement in a strict sense. An interesting and more recent development has been the use of general literacy tests administered to adults rather than to students. Since these tests provide a direct indicator of the basic skills and competences of the entire adult or working-age population, regardless of how these may have been acquired, in principle they are likely to be a better indicator of the general stock of human capital than measures of student achievement or competencies at a certain age or data on years of schooling.

Tables 1 to 3 summarize the most relevant achievement and literacy tests that have been administered to relatively broad samples of countries during the last several decades. The three tables have a common structure. For each test wave and subject, the table indicates the population being tested as characterized either by its age or by the school grade they are in and the number of countries in our reference sample of 21 OECD member states that participated in the test. The last two columns indicate the scoring scale being used (either percent correct or IRT, see below) and whether or not scores are directly comparable with those in more recent waves of the same test (DCLT). The relevant data are collected in an Excel file that is available with the paper.

a. Tests of student academic achievement

Table 1 lists the tests of student academic achievement that have been conducted by the International Association for the Evaluation of Educational Achievement (IEA). These tests seek to measure student achievement in three key areas (mathematics, science and reading) at three different stages in their education: primary school and lower and upper secondary school (4th and 8th grade and the final year of upper secondary education, FS). The last of these tests has a special version, known as TIMSS advanced (FSadv in the table), that is administered to students who are enrolled in advanced mathematics and physics programs or tracks.

The current generation of IAE's math and science tests goes under the name of TIMSS (for Trends in International Mathematics and Science Study). TIMSS started in 1995 (as the Third International Mathematics and Science Study) and has been administered every four years since then. As quite a few other international tests, TIMSS is scored using a methodology based on what is known as Item Response Theory (IRT) that takes into account the revealed difficulty of different test items. For the first edition of TIMSS, the grading scale was normalized to have a mean of 500 and a standard deviation of 100, which corresponded to the overall achievement distribution across all countries that participated in the test, assigning an equal weight to all of them. To allow TIMSS scores to be comparable over time within each subject, subsequent editions of the test retain a sufficient number of items from previous waves and the grading scale remains “constant”, not being renormalized to a mean of 500 each year.

The situation is the same for the current version of IAE's reading test, known as PIRLS (for Progress in International Reading Literacy Study), which is aimed at 4th grade students. PIRLS has been administered every five years starting in 2001. The grading scale is similar to the one used in TIMSS, with an original mean of 500 and standard deviation of 100 that correspond to the first edition and fully comparable scores for latter editions. Earlier versions of the IEA tests did not have a fixed periodicity but covered the same subjects as TIMSS and PIRLS at irregular intervals. With the exception of SIRS (Second International Reading Study), these tests reported scores simply as the percentage of correct answers.

We will work with a sample comprised by the initial OECD countries except for a few small and somewhat atypical economies such as Luxembourg and Iceland. These countries are, in particular, Australia, Austria, Belgium, Canada, Denmark, Finland, France, Germany, Greece, Ireland, Italy, Japan, Netherlands, New Zealand, Norway, Portugal, Spain, Sweden, Switzerland, United Kingdom and United States.
For a useful introduction to IRT proficiency scoring, see Annex B in OECD and Statistics Canada (2011).

Table 1: International tests of student achievement administered by the IEA

years of data collectionnamesubjectpopulation testedno. of countries in OECD21scaleDCLT
1964FIMSmath13 yrs, FS10, 10pcno
1970SRCreadingGrade 4, 8, FS7pcno
1970-71FISSscience10, 14, FS9, 11, 11pcno
1970-72FIRSreading13 yrs7pcno
1980-82SIMSmath8th, FS10, 8pcno
1983-84SISSscience5th, 9th, FS9,10, 9pcno
1990-91RLSreadingGrades 3-4, 7-817pcno
1990-91SIRSreading4th, 9th17, 17IRTno
1994-95TIMSS 95math4th, 8th, FS, FSadv11, 8, 13, 11IRTyes
1994-95TIMSS 95science4th, 8th, FS, FSadv10, 8, 13, 11IRTyes
1999TIMSS 99math8th grade7IRTyes
1999TIMSS 99science8th grade7IRTyes
2001PIRLS 01reading4th grade10IRTyes
2003TIMSS 03math4th & 8th grades10, 9IRTyes
2003TIMSS 03science4th & 8th grades9, 8IRTyes
2006PIRLS 06reading4th grade14IRTyes
2007TIMSS 07math4th & 8th grades13, 8IRTyes
2007TIMSS 07science4th & 8th grades12, 9IRTyes
2008TIMSS 08mathFSadv4IRTyes
2008TIMSS 08scienceFSadv4IRTyes
2011PIRLS 11reading4th grade14IRTyes
2011TIMSS 11math4th & 8th grades18, 10IRTyes
2011TIMSS 11science4th & 8th grades15, 10IRTyes
2015TIMSS 15math4th, 8th & FSadv18, 10, 6IRTyes
2015TIMSS 15science4th, 8th & FSadv18, 10, 6IRTyes
2016PIRLS 16reading4th grade18IRTyes
2019TIMSS 19math4th & 8th grades19, 13IRTyes
2019TIMSS 19science4th & 8th grades19, 13IRTyes

- Notes: final year of secondary schooling; advanced test, administered in the final year of secondary education. DCLT = directly comparable with latter tests pc = percent correct; IRT = scoring based on item response theory SRC = Study of Reading Comprehension; FIMS = First International Mathematics Study; FISS = First International Second International Science Study; RLS = reading literacy study; SIRS = Second International Reading Study. - Sources: Hanushek and Wößmann (2015), Altinok et al (2018) and IEA (https://www.iea.nl/studies).

In addition to those run by IEA, there have been a number of other international assessments of student performance. One of the most ambitious ones has been the MLA project (Monitoring Learning Achievement), organized by Unesco and Unicef, which covered over 70 countries, mostly LDCs (see Chinapah, 2003). There have also been a number of regional assessments. Many Latin American countries have joined UNESCO's Laboratorio Latinoamericano para la Evaluación de la Calidad de la Educación (LLECE) and many African countries participate in the South and Eastern African Consortium for Monitoring Educational Quality (SACMEQ) or the Programme d'Analyse des Systèmes Educatifs de la CONFEMEN (PASEC).

b. Tests of Student Literacy and Numeracy

Table 2 lists another family of tests, known as the PISA studies, that have been administered by the OECD every three years starting in 2000. PISA tests 15-year-old students just prior to the completion of mandatory schooling. The focus is not so much on academic achievement per se as on the command of the basic and applied skills that can be identified with a broad concept of literacy in the same three areas of interest as IEA tests (math, reading and science). In each wave of PISA, one of these three subjects is selected for a more in-depth analysis.

Table 2: International tests of student literacy administered by the OECD (PISA)

years of data collectionnamesubjectpopulation testedno. of countries in OECD21scaleDCLT
2000-02PISA 2000reading15 yrs19IRTyes
2000-02PISA 2000math15 yrs20IRTno
2000-02PISA 2000science15 yrs20IRTno
2003PISA 2003reading & math15 yrs20IRTyes
2003PISA 2003science15 yrs20IRTno
2006PISA 2006reading, math & science15 yrs20IRTyes
2009PISA 2009reading, math & science15 yrs20IRTyes
2012PISA 2012reading, math & science15 yrs21IRTyes
2015PISA 2015reading, math & science15 yrs21IRTyes
2018PISA 2018reading, math & science15 yrs20, 21, 21IRTyes

-Sources: OECD, https://www.oecd.org/pisa/

All PISA tests are scored using an IRT proficiency scale. Not all of them are directly comparable with latter studies, however. In particular, the reference scale for each subject has been set the first time the subject was the main focus of study. Hence, all reading tests are comparable because reading was the main subject of the first edition of PISA, but math scores are only directly comparable from 2003 onward and science scores from 2006 onward, after the first full assessment of each subject.

c. Tests of adult literacy

The final group of tests, listed in Table 3, are also literacy tests conducted by the OECD, but aimed now at the entire working-age population rather than at young people currently enrolled in school. Three successive studies (IALS, ALLS and PIAAC) have been conducted until now. All of them have tested reading and quantitative literacy while ALLS and PIACC also try to measure problem-solving abilities (not shown in the table). All three tests are scored using an IRT proficiency scale with a range from 0 to 500. As noted in OECD (2009), some results are comparable across tests. In particular, reading or literacy scores are directly comparable across all three of the surveys (after averaging prose and document literacy to obtain a single literacy score in the case of IALS and ALLS). Quantitative literacy scores from IALS are not directly comparable with numeracy scores from ALL and PIAAC (which are, however, directly comparable with each other) because the concept of numeracy used in the two more recent surveys is broader than the concept of quantitative literacy used in the earlier one. The last wave of PIAAC, finally, incorporates a section on ICT skills.

Table 3: International adult literacy tests administered by the OECD

years of data collectionnamesubjectpopulation testedno. of countries in OECD21scaleDCLT
1994-1998IALSreading literacy16-6515IRTyes
1994-1998IALSquantitative literacy16-6515IRTno
2003-08ALLSreading literacy16-658IRTyes
2003-08ALLSnumeracy16-658IRTyes
2013, 14-15PIACCreading literacy16-6517 + 2IRTyes
2013, 14-15PIACCnumeracy16-6517 + 2IRTyes
2014-15PIACCICT skills16-6514 +2IRT

- Key: IALS = International Adult Literacy Survey; ALLS = Adult Literacy and Lifeskills Survey; PIACC = Program for the International Assessment of Adult Competencies; final sec. = final year of upper secondary schooling. - Source: OECD, http://www.oecd.org/

While these data are of considerable interest because they provide the only available cross-country information on the skill level of the adult population, the short history of adult skill assessments is an important drawback. On the other hand, PIAAC sample sizes (over 5.000 per country) are sufficiently large to allow us to disaggregate the results by age group with some guarantee of representativeness and may therefore be used to construct synthetic time series of scores, as has been done by Coulombe and Tremblay (2006) using IALS data and by Schwerdt and Wiederhold (2018) with PIAAC.

3. Summary measures of schooling quality

Several groups of researchers have collected and homogenized the results of international student tests and have used them to construct summary performance measures that are usually interpreted as indicators of the quality of national educational systems or the level of skill the labor force. Among the most influential studies in this line of work are those of Eric Hanushek, Ludger Wößmann and different coauthors, and those of Nadir Altinok, Noam Angrist and other researchers linked to the World Bank.

Hanushek and Kimko (H&K 2000) construct an indicator of labor force quality for a sample of 31 (mostly advanced) countries using mean national scores in a number of international achievement tests in mathematics and science spread over several decades. To approximate the average quality of the labor force (rather than that of current students), H&K combine all the scores available for each country up until 1991 into a single cross-section indicator that is constructed as a weighted average of the normalized values of such scores (where the weights are based on the inverses of the country specific standard errors of the scores). They use two alternative normalization procedures to produce two different (but highly correlated) measures of labor force quality that they denote by QL1 and QL2. In the first case (QL1), the average world score in each year (measured by the percentage of correct answers) is normalized to 50. This procedure implicitly assumes that average performance does not vary over time. In the second case (QL2), they allow average performance to drift over time reflecting average US scores in a different but comparable set of national tests (NAEP).

The authors use the results of six such tests that were conducted between 1965 and 1991 (four by IEA and two by IAEP (International Assessment of Educational Progress).

Hanushek and Wößmann (H&W, 2012 and 2015) construct a refined version of QL2 for a sample of 64 countries (extended to 77 in the second study) using data from different tests of math and science conducted between 1964 and 2003 that include the first two waves of PISA. They standardize test scores prior to averaging them in order to put them all on the same distribution as the 2000 PISA test, with an overall mean of 500 for the OECD and an individual-level standard deviation of 100 for the same sample. In addition to average standardized scores (across assessments) for each country, they also report data on the average share of students that reach the thresholds for “basic” and “superior” performance, set at one standard deviation above and below the OECD average.

Standardized test scores are constructed as follows. First, the authors reconstruct the time path of absolute US performance starting from this country's results in PISA 2000 and going backward with the help of NAEP data. For each test conducted at time t, on subject s for age group a, an absolute normalized score for the US is calculated as

\[I _ {a s t} ^ {U S} = O _ {s, 2 0 0 0} ^ {U S, P I S A} + \frac {N A E P _ {a s t} ^ {U S} - N A E P _ {a s 1 9 9 9} ^ {U S}}{\overline {{S D}} _ {a s} ^ {U S , N A E P}} * S D _ {s, 2 0 0 0} ^ {U S, P I S A} \tag {1}\]

where is the original score of the US in PISA 2000 in subject s, NAEP is the age-, subject-, and time-specific NAEP test score, is the age- and subject specific standard deviation of the U.S NAEP test scores across individuals, calculated by averaging the available observations on standard deviations during the relevant period, and is the subject-specific standard deviation of U.S. students on the PISA 2000 test. Hence, changes in NAEP scores over time, relative to the 1999 edition of the test (the one closest to the PISA test used as a benchmark) are scaled up or down taking into account the difference in standard deviations of US individual scores across the two tests.

Next, other countries' normalized scores are calculated, taking into account their respective positions in relation to the United States. For each country , we have:

\[I _ {a s t} ^ {i} = I _ {a s t} ^ {U S} + \frac {O _ {a s t} ^ {i} - O _ {a s t} ^ {U S}}{S D _ {a s t} ^ {O S G}} * S D _ {s, 2 0 0 0} ^ {O S G, P I S A} \tag {2}\]

where denotes country i's original score for each subject and age group in the test conducted at time t. Differences in original scores between each country and the US are adjusted taking into account the cross-country variance in mean scores across a group of 13 advanced OECD countries (labeled by H&W as the OECD standardization group, OSG) that have participated in many of the relevant tests.

NAEP (National Assessment of Educational Progress) measures the performance of US students at different benchmark ages in around a dozen of subjects. One strand of the test (the so-called long-term trend assessments) provides nationally representative scores for math and reading for 9, 13 and 17-year old students at intervals of 2-4 years since the early 1970s that are measured on a consistent scale and can therefore be compared over time.
NAEP scores are available at 2-to-4-year intervals over the period; values for non-NAEP years are obtained by linear interpolation between available years.

In a series of papers, Nadir Altinok, Noam Angrist and various coauthors also construct and extend a database of standardized student achievement measures for a large number of countries. They augment H&W's sample by incorporating data on reading assessments that H&W's main indicator of cognitive skills disregards and information from other sources, such as regional achievement studies for countries that do not participate in global achievement studies and reading assessments. One of the latest versions of this database (AAP, 2018) provides data for 163 countries covering (unevenly) the period 1965-2015 at five-year intervals. Angrist et al (ADGP, 2021) add an additional country to the sample and extended the sample period until 2017.

Unlike H&K or H&W, AAP provide panel data with several observations for most countries that correspond to what they call harmonized learning outcomes (HLOs). HLOs are constructed as averages taken over different tests administered in the same or nearby years, after adjusting their results for differences in difficulty. In addition to mean scores (overall and disaggregated by educational level, subject, gender and other characteristics), they also report data on the percentage of students who reach three different benchmark levels (minimum, intermediate and advanced), thus providing useful information on the distribution of skills. Average country scores in different assessments at each point in time are standardized and brought into a common scale by using a procedure the authors refer to as pseudo-linear linking. A similar procedure is used to homogenize results over time, using NAEP data for the US as an anchor, as in H&W.

The standardization procedure essentially involves using the average scores obtained in each test by the set of countries that participate in both of them in order to calculate an “exchange rate” that can be used to adjust for differences in difficulty and grading scales. That is, given two tests X and Y, the score of country i in test X, , is converted to the scale of test Y using

\[y _ {i} = x _ {i} ^ {*} e \tag {3}\]

with

\[e = \frac {\mu (y)}{\mu (x)} = \frac {\frac {1}{n} \sum_ {i \in X \cap Y} y _ {i}}{\frac {1}{n} \sum_ {i \in X \cap Y} x _ {i}} \tag {4}\]

where the average scores for the two tests, and , are calculated over the n countries that have participated in both of them (i.e., over all ). When correcting for differences in difficulty over time, the exchange rate is based on US performance in international assessments and on NAEP.

In addition to the papers cited in the text, see also Altinok and Murseli (2007), Angrist, Patrinos and Schlotter (2013), Altinok, Diebolt and de Meulemeester (2014), Altinok and Angrist, Djankov, Goldberg and Patrinos (2019).
In the most recent version of this database (Angrist et al, 2021, supplementary information p.7) the standardization procedure is based on a regression of the form . The equation is estimated using

Filmer et al (2020) propose combining HLO's with data on average years of schooling to construct quality-adjusted or learning-adjusted years of schooling (LAYS). In particular, LAYS for country i are constructed as

\[{\mathrm{(5)}} L A Y S _ {i} = Y R S C H _ {i} \times Q _ {i} ^ {b}\]

where YRSCH is the average years of schooling of the population cohort and a measure of quality, relative to a benchmark level b. This benchmark may correspond to the top performing country or group of countries or to some other convenient reference level for good performance, e.g., a TIMSS score of 625 which corresponds to the threshold for advanced attainment set by TIMSS. While this measure depends in principle on the details of the grading scale, the specific test and subject chosen and the choice of benchmark level, the authors check that in practice the results do not change qualitatively with these factors. They report that correcting for quality increases cross-country differences, as countries where attainment levels are low typically also display poor performance in international assessments of student achievement.

Similarly, Kaarsen (2014) uses the variation in the results of achievement tests associated with an additional year of schooling in each country. On the other hand, Schoellman (2012) and Botev et al (2019) correct for quality using estimates of mincerian returns to schooling, i.e., the average wage increase linked to an additional year of education. The first of these studies uses data on immigrants to the US who have been educated in their country of origin, while the second one uses estimates of standard wage equations with data for the resident population, including migrants. An alternative proposal is the one used in the Penn World Table since its version 8 (see Feenstra, Inklaar and Timmer, 2015). The human capital index is computed using the average years of schooling (from Barro and Lee, 2013, Lee and Barro, 2001, Cohen and Leker, 2014, and de la Fuente and Doménech, 2006 and 2015) and an estimated rate of return to education, based on Mincer equation estimates around the world (from Psacharopoulos, 1994). An extension of this approach has been used by Angrist et al (2019), who also take into account an indicator of student learning or the quality of schooling.

4. Indicators of educational quality: potential uses and limitations in growth studies

The use of data from international achievement tests to construct indicators of educational quality or cognitive skills is certainly a relevant development that can help give us a better picture of cross-country stocks of human capital. Such data, however, have important limitations that should be kept in mind. An obvious one is the lack of quality indicators for tertiary education. Even more important is the relative scarcity of comparable cross-country data on student performance, as many countries have participated only in one or a few international assessments, mostly in recent years. As a result, there are only a few countries for which we have relatively long time series. The problem is compounded by the fact that most of these assessments measure the performance of students enrolled in primary or secondary schools, who have not yet entered the labor market. Hence, we are quite far from having the information that would be necessary to approximate, working with cohort data, the average quality of the human capital embodied in the labor force of most countries-- even for recent years, and much more so as we go back in time. An additional worry that arises when we try to go beyond the richest countries is that, certainly in past decades but even today, schooling is far from being universal in many countries, even at the lowest levels. For countries with low enrollment rates, the results of the student assessments described above measure the knowledge and competences of only a relatively small part of each cohort whose weight has likely been rising over time.

data for all countries that participate in both tests and is then used to estimate for those countries that have only participated in .
In principle, they would like to base the adjustment on a measure of how much is learned on average during an additional year in school, but this is difficult to estimate without simplifying assumptions that essentially bring us back to observed relative scores in achievement tests.
Hence, the indicator of quality-adjusted human capital would be of the form where r and are estimates of the returns to the quantity and quality of schooling.

The following calculations may help highlight the importance of the problems posed by the scarcity of student test data for their use in empirical studies of growth performance. Both PISA and IEA assessments are generally conducted with students between 10 and 16 years of age. Assuming the average individual remains in the labor force for 45 years, between ages 20 and 65, in order to approximate the average quality of the labor force in 2020, we would need test data for all those who entered the labor force during the previous 45 years, i.e., between 1975 and 2020, who were tested between 5 and 10 years earlier, that is, between 1965 and 1970. Since the earliest assessments we have were conducted in those years (in a handful of countries), all the test data we have accumulated to date would only allow us to approximate the skill level of today's labor force in a few countries – but certainly not its average quality over the last several decades, which is the variable that should be included in many of the growth regressions that have been run in the literature.

Hence, the available data on student performance is clearly insufficient to construct time series of stock measures of average skill for the labor force, but they can still be quite useful as a flow measure of investment in quality at each point in time. To exploit these data in growth studies, we need to use empirical specifications that are suitable for flow data. One possibility that has been used in the literature to get around similar problems regarding other growth determinants is the specification developed by Mankiw, Romer and Weil (MRW, 1992) as a log-linear approximation around the steady state of a generalized Solow model. This approach may be particularly useful in combination with pooled data at relatively high frequencies as a way to exploit the time variation in the data.

Another important limitation of using student test data in empirical growth equations is the high potential for severe endogeneity and reverse causation problems. Economic growth generates increased public and private resources that may be used by governments and families to finance higher quality educational systems that yield higher student performance, as well as more years of schooling. If we focus on flow measures such as test scores (or enrollment rates), the feedback effect from growth to education can be quite rapid, while higher test scores will only affect growth much further into the future, when today's students enter the labor market and become employed. As a result, reverse causation is likely to be an important problem even in data at relatively high frequencies. On the other hand, when we rely on data on average years of schooling of the adult population, the direct effect of schooling on growth should be immediate, while feedback effects from growth to increased average schooling will involve much longer lags. As a result, average schooling levels can be considered as a predetermined variable except over rather long periods, and reverse causation should be much less of a problem.

The problems discussed above do not arise in literacy assessments of the entire adult population, such as PIAAC, but in this case we only have very recent results for relatively few countries. As we have seen, a possible way to mitigate the problem this represents is to construct synthetic time series of adult competences using the age distribution of the microdata of these tests, as has been done by Coulombe and Tremblay (2006) using IALS data and Schwerdt and Wiederhold (2018) with PIAAC. In this regard, Castelló (2018) has shown that the largest effects on economic growth are linked to the human capital stock of the population aged 40–49 years for a large sample of 146 countries. That is, the population in the middle of its professional career is the most representative cohort of the working-age population in terms of productivity. This result implies that the most relevant variable for explaining today's growth is the quality of schooling between 30 and 35 years ago, when this central cohort was in primary and secondary school.

Synthetic time series, however, raise several complications that have to do with the fact that the scores of the different cohorts are measured at a single point in time which corresponds to a different age for each of them. This may introduce a bias if, for instance, skills change significantly over time (due for instance to the accumulation of experience and then to aging and depreciation), or if survival or migration rates are correlated with skill levels, as seems likely.

More generally, it is clear that none of the existing assessments can cover and measure correctly all the competences and skills that determine the productivity of a country's labor force or its capacity to innovate, including those acquired in universities and workplaces or the specialized and highly complex knowledge and skills of scientists and high-level technicians that go well beyond what these tests measure. For all these reasons, it is important to exercise caution when using and interpreting existing indicators of educational quality and of the skill level of the overall population. In countries with high enrolment rates where we have relatively long series of test results and these remain stable over time, we may perhaps be somewhat confident that student achievement data can give us some idea of the average quality of the school system, which is certainly an important input for growth, but not the only one. In countries with low enrolment rates, shorter series or where test results have changed significantly over time, we have to be even more cautious. As for adult literacy data, we must keep in mind that they pick up only a (possibly small) fraction of the relevant knowledge and skills.

Finally, it seems obvious that both the quantity and the quality of education must be taken into account in order to correctly approximate the stock of human capital, for both are essential inputs in its production. Keeping a large fraction of the population in school for many years may be a bad investment if the poor quality of education prevents students from acquiring the skills the productive system demands, but an excellent educational system will have only a limited effect on productivity if it excludes most of the population. Unless we have good direct measures of the relevant knowledge and skills of the entire population --which we surely do not, except possibly for very recent years-- we need to measure as well as possible both dimensions of the stock of educational capital and use them jointly in empirical analysis. A promising possibility in this line consists in adjusting years of schooling for quality, as has been done in some recent studies using alternative procedures (see for instance Filmer et al, 2020 and Angrist et al, 2019), although the scarcity of data implies that stock quality measures will generally have no time variation, which will limit their usefulness in empirical analyses and other applications.

5. A quick look at the data

Table 4 collects some educational indicators of interest, mostly referring to years around 2010, for a sample of 21 OECD countries. All variables are normalized, with their unweighted cross-country averages set to 100. Column [1] shows an indicator of adult skills, constructed as the average of literacy and numeracy scores. For most countries, the data come from PIAAC. For those that did not participate in this assessment, we use the results of the most recent similar test that is available, i.e. ALL in the case of Switzerland and IALS in that of Portugal. Column [2] shows average PISA scores in 2012 and column [3] average scores (across available subjects and grades in each country) in the 2011 round of IEA tests (TIMSS and PIRLS). The following three columns contain summary indicators based on student achievement and skills tests. Average harmonized learning outcomes (HLOs) in 2010, taken from AAP (2018) and shown in column [4], contain information on both student achievement and student skills. Column [5] shows H&W's (2015) indicator of population cognitive skills, which is constructed by averaging student achievement and literacy tests on math and science (but not reading) over several decades. Column [6] contains a similar “stock” indicator of educational quality, the cumulative average of available HLO's until 2005, which also incorporates reading results. Column [7] shows our estimate of average years of schooling in 2010, taken from de la Fuente and Doménech (D&D, 2015) and column [8] relative real GDP per working-age person (relative income per capita, for short, from now on), taken from an updated version of the data set used in D&D (2006). Finally, column [9] shows the average value between 2005 and 2015 of the measure of social welfare proposed by Jones and Klenow (2016), which is computed by aggregating (with the appropriate utility weights) private and public consumption per capita, and indicator of income equality, hours worked and life expectancy.

Table 5 displays pairwise correlations for the variables shown in Table 4. It should be noted that correlations between quality variables, while always positive, are often fairly low, suggesting that it may be difficult to construct a single indicator that adequately summarizes educational quality. We can classify the educational indicators we have gathered into two groups: flow measures of student performance (measuring the academic achievement or basic skills of each young cohort), which can be seen as indicators of the quality of education at a given point in time, and stock measures of the quantity or quality of schooling for the entire adult or working-age population, which are sometimes constructed by averaging flow indicators over long periods. Correlations tend to be higher within each of these groups than across them, although with some exceptions. As should be expected, stock measures are more highly correlated with income per capita than flow measures. As a summary indicator of student performance, we will use AAP's HLOs, which combine information on all PISA and IAE scores and display a fairly high correlation with both of these variables (0.743 and 0.812, respectively). As for the stock measures, we will focus on adult skills and years of schooling as direct measures of quality and quantity, and retain for some purposes the other two variables, H&W's indicator of cognitive skills and the average value of available HLOs until 2005.

This data set has been compiled using OECD data on member states' national accounts, working-age populations and a set of OECD-specific purchasing power parities.

Table 4: Selected normalized education and income indicators around 2010

[1][2][3][4][5][6][7][8][9]
adult skills2012 student skills2011 student achieve2010 HLOH&WHLO avge till 20052010 years of schooling2010 ypc 15-642005-15 avge welfare
Australia102.1101.698.498.7102.595.3105.9110.7109.1
Austria101.499.2100.199.2102.4105.4101.4105.1112.1
Belgium103.5101.099.8103.4101.5107.095.9101.1103.9
Canada100.4103.5100.8100.4101.4101.4112.9103.5104.5
Denmark102.398.7103.3101.999.9106.2103.199.8101.7
Finland106.1104.9101.4102.1103.2100.7102.696.9102.5
France96.299.099.6100.0101.498.0101.092.2106.4
Germany100.8102.2101.9102.499.795.4103.8102.0103.0
Greece94.292.396.593.392.794.886.072.969.1
Ireland97.2102.299.9100.3100.5102.998.5106.284.1
Italy92.797.198.596.995.894.184.983.498.0
Japan108.8107.1108.8111.4106.9112.8105.698.588.5
Netherlands105.1102.8103.2104.7102.997.9105.1110.8109.7
New Zealand102.8101.096.694.6100.291.596.176.977.7
Norway103.798.394.493.097.298.3111.4144.3124.9
Portugal84.396.7100.999.791.990.672.266.865.7
Spain92.797.095.394.997.298.981.980.089.6
Sweden104.095.598.595.7100.995.1113.9106.3117.8
Switzerland105.1102.797.8106.9103.5110.6104.9114.1122.0
United Kingdom99.599.6102.099.499.6104.898.5100.498.4
United States97.397.5102.4101.398.798.4114.3128.0111.3
average100.0100.0100.0100.0100.0100.0100.0100.0100.0

Table 5: Correlations between pairs of indicators

correlation with:student skillsstudent achieveHLOH&WHLO avge till 2005years of schoolingypc 15-64Welfare
adult skills (PIAAC)0.6410.2950.4480.8390.5450.7540.5590.570
student skills (PISA)1.0000.5480.7430.8120.5250.4070.2650.206
student achievement (IAE)1.0000.8120.4810.4630.2410.047-0.049
HLO1.0000.6690.6860.2700.1770.160
H&W1.0000.6400.6580.4180.502
HLO avge until 20051.0000.3290.3390.310
years of schooling1.0000.8160.755
ypc 15-640.826

Figure 1: Selected human capital indicators around 2010 Unweighted sample average = 100

Figure 1: Selected human capital indicators around 2010
Unweighted sample average = 100

Figure 1 displays the cross-section profile of the three main indicators we have selected, those that measure adult skills, student performance, and average years of schooling of the adult population, with countries ordered by the adult skills indicator. In terms of this variable, Southern European countries display the lowest scores, followed by the Anglo-Saxon and Central European nations, while Northern Europe and Japan perform best. Country performance in terms of student competences and average years of schooling, however, often deviates markedly from this pattern. For instance, the US, Canada, Norway and Sweden do much better in terms of years of schooling than in adult skills, while the opposite is true in Southern Europe. Roughly speaking, student performance measures tend to lie above adult skills for lower values of the latter variable and below them in the upper half of the distribution.

Table 6 shows country rankings according to the same three indicators, together with each country's average rank and its rank range, defined as the difference between its highest and lowest rankings. Looking at the table, it is clear that the three indicators generate rather different rankings. In some cases, the differences across indicators for a given country are quite striking. For instance, Sweden and Norway do quite well in terms of adult skills (where they rank in positions 5 and 6) but very poorly in terms of student performance in standardized tests (where they drop to positions 17 and 21 respectively), and the US goes from the first position in terms of years of schooling to the when we consider adult skills. Japan, the Netherlands, Switzerland and Finland are well ranked in terms of adult and student performance, but not so much when it comes to years of schooling, and Southern Europe displays consistently poor performance in terms of all indicators, with the partial exception of Portugal in the case of student performance.

Table 6: Country rankings around 2010

adult skillsstudent perf. HLOsyears of schoolingaverage rankrange max – min rank
Japan1162.75
Netherlands3374.34
Switzerland4284.76
Finland26116.39
Sweden51728.015
United States15818.014
Canada13938.310
Denmark97108.73
Germany12598.77
Belgium74179.313
Australia1015510.010
Norway621410.317
Austria11141212.33
France17111313.76
Ireland16101513.76
United Kingdom14131413.71
New Zealand8191614.311
Italy19161918.03
Portugal21122118.09
Greece18201818.72
Spain20182019.32

As we have already indicated, we expect that both years of schooling and educational quality should contribute positively to adult skills. As a very rough test of this hypothesis, we can regress the PIAAC-based indicator of adult skills, which a priori would seem to be the best available proxy for this variable, on years of schooling (yrsch) and one of the two stock quality indicators we have selected (h&w and hlo_at05), with all variables measured in logs.

As can be seen in Table 7, the results are consistent with our hypothesis that both quantity and quality matter. Years of schooling and educational quality are always significant, whether entered alone or jointly in the equation. H&W's science and math-based cognitive skills indicator performs better than the cumulative average of HLOs, but in both cases the strategy of averaging flow performance measures over several decades seems to be successful at producing a stock indicator of quality that helps explain average levels of adult skills. Incidentally, the high R-squared of these regressions suggest that we are likely to run into severe multicollinearity problems if we try to use several educational indicators as explanatory variables for income or welfare levels or growth rates in the same equation. One way to mitigate this problem may be to use quality-adjusted years of schooling. With the variables in logs, this basically involves adding up the quantity and quantity variables to leave a single regressor, a procedure that would only be justified if we cannot reject the hypothesis that the coefficients of the two variables in the regression are equal. Looking at equations [2] and [3], the relevant coefficients are similar in the case of hlo, but not when we use h&w as a quality indicator, suggesting that quality adjustment would be acceptable for the first variable but not for the second. In equation [6] we impose the equality restriction and check that HLO-based quality-adjusted years of schooling (qayrsch) performs rather well – although slightly less so than lh&w alone.

Table 7: Determinants of adult skills around 2010

[1][2][3][4][5][6]
lyrsch0.397(5.66)0.343(4.95)0.201(2.69)
lhlo_at050.290(2.14)0.536(2.85)
lh&W0.920(3.79)1.374(6.81)
$lqayrsch^*$ 0.328(6.67)
$R^2$ 0.62750.70310.79270.70920.29910.7010

- Note: the dependent variable, , is the log of the adult skills indicator. All equations include a constant that is not reported. All regressors are measured in logs. (*) Quality adjustment based on hlo_at05, taking as a benchmark the sample average of this indicator.

Table 8 displays the results of separate regressions of individual PIAAC scores on years of schooling, time elapsed since the completion of schooling and the square of this last variable for those countries for which all the required data are available. The correlation between years of schooling and PIAAC scores is very strong and statistically significant in all countries, suggesting that, as may be expected, skills are gradually acquired over time in school. The expected contribution of a year of schooling to the average PIAAC score ranges between 4,77 points in Italy to 8,36 points in Germany, but it is not clear that the value of this coefficient can be interpreted as an indicator of school quality, as intercept coefficients also vary widely within the sample and do so in a way that tends to offset slope differences. Figure 2a shows the estimated effect of years of schooling on adult skills. On average, each additional year of schooling increases PIAAC scores by approximately 6 points.

Except for Greece and the UK, PIAAC scores are lower for individuals who left school a long time ago than for more recent graduates with the same level of schooling as shown in Figure 2b. This pattern suggests that school-acquired knowledge and competences depreciate over time, possibly as a result of aging and obsolescence, but may also reflect an increase in the quality of schooling over time. There are, however, some exceptions and significant differences across countries in the rate at which skills seem to depreciate over time. In any case, the effect of time elapsed since graduation is relatively small: on average adult skills scores are 20 points lower after 45 years since graduation.

Table 8: PIAAC average score in math and reading as a function of years of schooling and time elapsed since completion of studies individual data by country

yrs schoolTime since gradTime since grad sqconstantN obsRsq
Germany8,36(37,48)-0,91(6,09)0,00(0,16)176,70(54,64)5.0880,322
Belgium7,39(35,20)-0,85(6,33)0,00(1,16)203,50(71,25)4.9100,324
Denmark6,85(32,75)-0,46(3,54)-0,01(1,74)200,00(73,29)7.1670,234
Spain6,20(38,48)-0,45(3,68)0,00(1,06)191,90(82,93)5.6890,333
Finland5,94(23,11)-0,92(5,76)0,00(0,80)227,70(69,70)5.4200,259
France7,30(44,49)-0,88(7,49)0,01(3,72)189,40(82,28)6.6170,355
Ireland6,53(27,10)-0,33(2,19)0,00(0,86)171,30(41,95)5.8900,239
Italy4,77(22,82)-0,38(2,34)0,00(0,40)207,00(57,29)4.5060,237
Japan6,53(31,57)+0,46(3,75)-0,02(8,83)213,80(72,38)5.1470,301
Norway7,05(23,43)-0,16(0,99)-0,01(2,14)186,70(44,84)4.8880,196
Netherlands6,50(24,71)-0,34(2,33)-0,01(3,04)210,00(55,90)4.9680,269
UK7,53(23,10)+0,56(3,04)-0,01(2,38)170,20(37,40)7.5490,145
Sweden7,95(25,80)-0,87(5,16)0,01(2,43)195,00(49,74)4.3630,207

- Notes: - t statistics in parentheses below estimated coefficients - The dependent variable is the individual's PIAAC score, measured as the average value of the math and reading scores. The regressors are the number of years of schooling, the time elapsed since the completion of schooling and the square of this last variable.

Figure 2: Predicted PIAAC score by country as a function of

Figure 2: Predicted PIAAC score by country as a function of

b. time elapsed since graduation

Figura

Switching from the cross-section to the time-series dimension, we are interested in the stability of flow measures of educational performance over time. The assumption that quality levels do not change much over time has been made in many studies (see, for example, H&K) as a convenient way to get around limited data availability, for it allows us to approximate the skill level of the entire adult population using data on the educational performance of current and recent cohorts of students. It is not clear, however, that this is indeed the case.

As shown in Figure 3, a first look at HLO scores, the indicator with more observations in our sample, shows a lot of variation over time, some strange patterns and positive trends for many countries and for the sample average. Thus, the average score for the 21 countries has increased steadily, except in 1980, between 1970 and 2015, rising from 467 to 522, with a 11,6% increase in educational performance. The improvement in scores has been significant in the case of Portugal: in 1990 was the country with the worst performance in the sample, with a HLO score of 397, but 25 years later was the 7 country with better performance, between the US and Germany, having registered a 32% increase in scores. At the same time, there are some surprising observations as, for example, Germany in 1975 or Australia, Finland and France in 1980.

Figure 3: HLO scores over time, 1970-2015, 21 OECD countries

Figure 3: HLO scores over time, 1970-2015, 21 OECD countries

To gauge the degree of stability of country performance over time, we estimate country-specific trends as follows. Given an educational indicator, x, let

\[(7) \Delta x _ {n} = \frac {x _ {n} - x _ {n - 1}}{t _ {n} - t _ {n - 1}}\]

be its average annual variation between observations n-1 and n, dated at and respectively. For each country, we estimate a regression of the form

\[(8) \Delta \mathbf {x} _ {n} = g ^ {*} \mathbf {x} _ {n - 1} + \varepsilon\]

where the constant has been suppressed and is a random disturbance.

Table 9: Estimated country specific trends of some indicators of educational quality

student skills (PISA)Student achievement (IAE)HLOs
g(t)g(t)g(t)
Australia-0.33%(4.25)+0.12%(2.23)-0.13%(0.12)
Austria-0.23%(1.07)+0.04%(0.13)-0.10%(0.65)
Belgium-0.08%(0.64)-0.06%(0.13)-0.03%(0.28)
Canada-0.16%(1.55)+0.16%(0.69)+0.24%(2.00)
Denmark+0.04%(0.29)+0.13%(0.38)+0.03%(0.21)
Finland-0.24%(1.13)-0.44%(0.72)-0.11%(0.14)
France-0.15%(0.94)-0.42%(2.02)+0.16%(0.22)
Germany+0.14%(0.45)+0.16%(0.42)+0.50%(1.18)
Greece-0.09%(0.44)--+0.37%(0.72)
Ireland-0.11%(0.40)+0.16%(0.45)+0.42%(1.12)
Italy+0.03%(0.11)-0.20%(0.78)+0.35%(0.98)
Japan-0.25%(0.80)+0.12%(0.99)+0.17%(2.28)
Netherlands-0.29%(2.51)+0.04%(0.11)+0.27%(1.27)
New Zealand-0.31%(1.90)-0.09%(0.30)+0.28%(0.69)
Norway-0.06%(0.23)+0.12%(0.23)+0.13%(0.10)
Portugal+0.34%(0.95)0.00%(0.01)+1.05%(3.18)
Spain-0.06%(0.28)+0.09%(0.29)+0.07%(0.20)
Sweden-0.12%(0.45)-0.21%(0.45)+0.34%(0.86)
Switzerland-0.09%(0.48)-0.09%(0.87)
United Kingdom-0.17%(0.82)+0.15%(0.86)-0.14%(0.32)
United States-0.05%(0.18)+0.04%(0.36)+0.26%(3.85)
Corr. con PISA1.0000.1240.665
Max no. of obs.669

The results, shown in Table 9, suggest that educational quality is not stable over time in many countries. At the standard confidence level of 95%, only between two and four countries display trends that are significantly different from zero depending on the specific variable we examine. On the other hand, this test may be too stringent given the small number of observations, which go from a minimum of 6 to a maximum of 9 per country, depending on the indicator. If we consider all estimates with a t ratio of roughly one or greater, the number of countries for which there are fairly clear indications of a positive or negative trend rises sharply, raising increasing doubts on the validity of the constant quality assumption that has often been used in the literature to justify the use as an explanatory variable for growth of average test scores computed over different time periods depending on data availability on each country. This result, however, does not raise doubts about the potential usefulness of such data in combination with more appropriate flow specifications.

6. Conclusion

Scores in standardized international student achievement tests and some recent adult literacy studies provide interesting data on the quality of educational outputs and on the skill level of the population that can be a useful complement to the data on the quantity of schooling that have been most commonly used in the growth literature. The use of these data is likely to improve our ability to measure human capital accurately and help us understand its contribution to output levels, economic growth and social welfare. In this paper we have reviewed the main sources of primary data in this area and some recent efforts to systematize and standardize them with a view to constructing useful summary indicators of the quality of educational systems or the skill level of the adult population. We have used these data to look at cross-country educational performance in recent years within a sample of core OECD countries and discussed some of their limitations in terms of their potential use in empirical growth studies.

Accepting that quality matters in education does not mean that quantity should be ignored. The skill level of the labor force will surely depend on both the quantity and the quality of schooling. We have provided some preliminary evidence in favor of this view and argued that progress in this area is most likely to come from studies that try to combine both dimensions.

The scarcity of quality data, both across countries and over time, however, will be a serious handicap in this effort. We do not have long enough series on student performance to construct good stock measures of quality by averaging scores over a sufficiently long period, but we may be able to get around this difficulty by using flow data on quality in MRW-type panel specifications. Another possibility may be to try to correct years of schooling for quality before using them to estimate standard growth or productivity equations with panel data, although the quality indicators required for the correction are likely to lack time variation. A third route relies on the construction of synthetic time series of adult skill indicators using the available information on the age distribution of PIAAC results, and possibly correcting for estimated depreciation over time. This approach is likely to become more productive as new waves of PIAAC become available.

References

  1. Mullis, I. V. S., Martin, M. O., Foy, P., Kelly, D. & B. Fishbein. (2020). TIMSS 2019 International Results in Mathematics and Science. Retrieved from Boston College, TIMSS & PIRLS International Study Center website: https://timssandpirls.bc.edu/timss2019/international-results/download-center/

Data Appendix

Data on scores on IAE tests

Results for TIMSS 2019 are taken from

References

  1. Results for TIMSS 2015 are taken from

References

  1. Mullis, I. V. S., Martin, M. O., Foy, P., & Hooper, M. (2016). TIMSS 2015 International Results in Mathematics. Retrieved from Boston College, TIMSS & PIRLS International Study Center website: http://timssandpirls.bc.edu/timss2015/international-results/ or http://timssandpirls.bc.edu/timss2015/international-results/download-center/

References

  1. Data on earlier TIMSS tests are from exhibits 1.1, 1.2, 1.5 and 1.6 of

References

  1. Mullis, I., M. Martin, P. Foy and A. Arora (2012). TIMSS 2011 International Results in Mathematics. TIMSSS & PIRLS International Study Center, Lynch School of Education, Boston College and International Association for the Evaluation of Educational Achievement, Amsterdam. http://timssandpirls.bc.edu/timss2011/international-results-mathematics.html

References

  1. The only exception has to do with the results of TIMSS95 for the final year of secondary schooling, which come from Tables 2.1 and 2.2 of

Mullis, I., M. Martin, A. Beaton, E. Gonzalez, D. Kelly and T. Smith (1998). Mathematics and science achievement in the final year of secondary school. IEA's Third International Mathematics and Science Study (TIMSS). TIMSS International Study Center, Boston College. http://timssandpirls.bc.edu/timss1995i/MathScienceC.html

References

  1. PIRLS data are taken from exhibits 1.1, and 1.5 of

Mullis, I., M. Martin, P. Foy and K. Drucker (2012). PIRLS 2011 International Results in Reading. TIMSSS & PIRLS International Study Center, Lynch School of Education, Boston College and International Association for the Evaluation of Educational Achievement, Amsterdam. http://timssandpirls.bc.edu/pirls2011/international-results-pirls.html

References

  1. The results of earlier tests have been taken from:

Lee, J.-W. and R. Barro (2001). "Schooling Quality in a Cross-Section of Countries." Economica vol. 68, no. 272, pp. 465-88.

who report results on a percentage-correct basis.

Referenced IAE publications are available online at:

http://www.iea.nl/completed_studies.html

Data on PISA scores

Most of the data are taken from the report on PISA 2012:

OECD (2014). PISA 2012 Results in Focus. What 15-year-olds know and what they can do with what they know. Programme for International Student Assessment. Paris. http://www.oecd.org/pisa/keyfindings/pisa-2012-results.htm

See in particular Annex B1, tables 1.2.3b, 1.4.3b and 1.5.3b

the rest of the data (on math and science for 2000 and 2003) are taken from:

PISA 2018: http://www.oecd.org/pisa/PISA-results_ENGLISH.png

PISA 2015: http://www.oecd.org/pisa/pisa-2015-results-in-focus.pdf

OECD (2004). Learning for Tomorrow's World. First results from PISA 2003. Programme for International Student Assessment. Paris. http://www.oecd.org/edu/school/programmeforinternationalstudentassessmentpisa/learningfortomorrowsworld-englishversion-chapterbychapter.htm

Annex B1, tables 6.6 and 2.5c

OECD and Unesco Institute for Statistics (2003). Literacy Skills for the World of Tomorrow. Further results from PISA 2000. Programme for International Student Assessment. Paris. http://www.keepeek.com/Digital-Asset-Management/oecd/education/literacy-skills-for-the-world-of-tomorrow_9789264102873-en#page286

Annex B1, tables 3.1 and 3.2

The relevant publications are available at http://www.oecd.org/pisa/keyfindings/

Data on Adult Literacy Surveys

IALS final report, table 2.1 pp. 135-6

OECD and Statistics Canada (2000). Literacy in the information age. Final report of the International Adult Literacy Survey. Paris. http://www.oecd.org/edu/skills-beyond-school/41529765.pdf

ALL, table 2.2 in

OECD and Statistics Canada (2011). Literacy for life: Further results from the Adult Literacy and Life Skills Survey. Second International ALL Report. OECD Publishing. http://www.statcan.gc.ca/pub/89-604-x/89-604-x2011001-eng.pdf

There are some data by age group in Table 2.6.1 in p. 65

PIAAC

OECD (2013). OECD Skills Outlook 2013. First Results from the Survey of Adult Skills. OECD Publishing, Paris. http://skills.oecd.org/OECD_Skills_Outlook_2013.pdf

Data for the 16-65 population from Tables A2.2a and A2.6a, data broken down by age group from tables A3.2

AAP (2018) Harmonized learning outcomes database

Altinok, N., N. Angrist and H. Patrinos (AAP, 2018). “Global Data Set on Education Quality (1965–2015).” World Bank, Policy Research Working Paper no. 8314.

Data downloaded from

https://github.com/owid/owid-datasets/tree/master/datasets/Global%20Data%20Set%20on%20Education%20Quality%20(1965-2015)%20-%20Altinok%2C%20Angrist%2C%20and%20Patrinos%20(2018)

References

  1. Altinok, N. (2007). “Human capital quality and economic growth.” IREDU Working Paper no. 2007/1, Université de Bourgogne. http://bit.ly/3c8frlf
  2. Altinok, N., N. Angrist and H. Patrinos (AAP, 2018). “Global Data Set on Education Quality (1965–2015).” World Bank, Policy Research Working Paper no. 8314.
  3. Altinok, N., and Murseli, H. (2007). "International database on Human Capital Quality." Economics Letters 96(2), pp. 237-244.
  4. Altinok, N., Diebolt, C., & de Meulemeester, J.-L. (2014). "A New International Database on Education Quality: 1960-2010." Applied Economics 46 (11), 1212-1247.
  5. Angrist, N., S. Djankov, P. Goldberg and H. Patrinos (2019). "Measuring Human Capital." Policy Research Working Paper 8742, World Bank, Washington DC.
  6. Angrist, N., S. Djankov, P. Goldberg and H. Patrinos (2021). “Measuring human capital using global learning data.” Nature. https://doi.org/10.1038/s41586-021-03323-7
  7. Angrist, N., Patrinos, H.A., & Schlotter, M. (2013). “An Expansion of a Global Data Set on Educational Quality.” (No. WPS6536). World Bank Policy Research Working Paper no. 6536.
  8. Barro, R. J. and J.W. Lee (2013), “A new data set of educational attainment in the world, 1950-2010” Journal of Development Economics 104: 184–198.
  9. Botev, J., B. Égert, Z. Smidova y D. Turner (2019). “A new macroeconomic measure of human capital with strong empirical links to productivity.” OECD, Economics Department working paper no. 1575, Paris. https://dx.doi.org/10.1787/d12d7305-en
  10. Breton, T. (2011). “The quality vs. the quantity of schooling: What drives economic growth?” Economics of Education Review, 30, pp, 765-73.
  11. Castelló-Climent, A. (2019). “The age structure of human capital and economic growth.” Oxford Bulletin of Economics and Statistics, 81(2), 394-411.
  12. Chinapah, V. (2003). "Monitoring Learning Achievement (MLA) Project in Africa", ADEA Biennial Meeting 2003. Grand Baie, Mauritius, 3-6 December. http://www.adeanet.org/adea/biennial2003/papers/2Ac_MLA_ENG_final.pdf
  13. Cohen, D. and L. Leker (2014). "Health and Education: Another Look with the Proper Data", mimeo Paris School of Economics.
  14. Coulombe, S. and J. F. Tremblay (2006). "Literacy and Growth," B.E. Journals in Macroeconomics: Topics in Macroeconomics Vol. 6, No. 2, pp. 1-32.
  15. de la Fuente, A. and R. Doménech (D&D, 2006). “Human capital in growth regression: How much difference does data quality make?” Journal of the European Economic Association 4(1), pp. 1–36.
  16. de la Fuente, A. and R. Doménech (D&D, 2015). “Educational Attainment in the OECD, 1960-2010. Updated series and a comparison with other sources.” Economics of Education Review 48, October, pp. 56–74.
  17. Feenstra, R., R. Inklaar and M. Timmer (2015). “The Next Generation of the Penn World Table.” American Economic Review, 105(10), 3150-3182. https://bit.ly/3udI5rH
  18. Filmer, D., Rogers, H., Angrist, N., & Sabarwal, S. (2020). “Learning-adjusted years of schooling (LAYS): Defining a new macro measure of education.” Economics of Education Review, 77, 101971.
  19. Folloni, G., and Vittadini, G. (2010). “Human capital measurement: a survey.” Journal of Economic surveys, 24(2), 248-279.
  20. Hanushek, E. and D. Kimko (2000). "Schooling, labor-force quality and the growth of nations." American Economic Review 90(5), pp. 1184-208.
  21. Hanushek, E. and L. Wößmann (H&W, 2010). “The economics of international differences in educational achievement.” NBER Working Paper no. 15949, Cambridge, Mass.
  22. Hanushek, E. and L. Wößmann (2012). “Do better schools lead to more growth? Cognitive skills, economic outcomes and causation.” Journal of Economic Growth 17, pp. 267-321.
  23. Hanushek, E. and L. Wößmann (2015). The Knowledge Capital of Nations: Education and the Economics of Growth. CESifo book series, MIT Press. Cambridge, Mass.
  24. Heston, A., R. Summers, and B. Aten (2002). "Penn World Table Version 6.1." Center for International Comparisons of Production, Income and Prices at the University of Pennsylvania.
  25. Institute of Education Sciences (IES, 2010). An introduction to NAEP. http://nces.ed.gov/nationsreportcard/about/
  26. Jones, C. I., and Klenow, P. J. (2016). "Beyond GDP? Welfare across countries and time." American Economic Review, 106(9), 2426-57.
  27. Kaarsen, N. (2014). “Cross-country differences in the quality of schooling.” Journal of Development Economics, 107, 215-224.
  28. Lee, J.-W. and R. Barro (2001). "Schooling Quality in a Cross-Section of Countries." Economica vol. 68, no. 272i, pp. 465-88.
  29. Mankiw, N. G., Romer, D., and Weil, D. N. (1992). “A contribution to the empirics of economic growth.” The Quarterly Journal of Economics, 107(2), 407-437.
  30. Mullis, I., M. Martin, P. Foy and A. Arora (2012). TIMSS 2011 International Results in Mathematics. TIMSSS & PIRLS International Study Center, Lynch School of Education, Boston College and International Association for the Evaluation of Educational Achievement, Amsterdam. http://timssandpirls.bc.edu/timss2011/international-results-mathematics.html
  31. Martin, M., I. Mullis, P. Foy and G. Stanco (2012). TIMSS 2011 International Results in Science. TIMSSS & PIRLS International Study Center, Lynch School of Education, Boston College and International Association for the Evaluation of Educational Achievement, Amsterdam. http://timssandpirls.bc.edu/timss2011/international-results-science.html
  32. Mullis, I., M. Martin, P. Foy and K. Drucker (2012). PIRLS 2011 International Results in Reading. TIMSSS & PIRLS International Study Center, Lynch School of Education, Boston College and International Association for the Evaluation of Educational Achievement, Amsterdam. http://timssandpirls.bc.edu/pirls2011/international-results-pirls.html
  33. Mullis, I., M. Martin, A. Beaton, E. Gonzalez, D. Kelly and T. Smith (1998). Mathematics and science achievement in the final year of secondary school. IEA's Third International Mathematics and Science Study (TIMSS). TIMSS International Study Center, Boston College. http://timssandpirls.bc.edu/timss1995i/MathScienceC.html
  34. National Center for Education Statistics (NCES, 2013). NAEP 2012. Trends in Academic Progress. Reading 1971-2012, Mathematics 1973-2012. Institute of Education Sciences. http://nces.ed.gov/pubsearch/pubsinfo.asp?pubid=2013456
  35. OECD (2004). Learning for Tomorrow's World. First results from PISA 2003. Programme for International Student Assessment. Paris. http://www.oecd.org/edu/school/programmeforinternationalstudentassessmentpisa/learningfortomorrowsworld-englishversion-chapterbychapter.htm
  36. OECD (2009). “International adult literacy and basic skills surveys in the OECD region.” Education Directorate, Working Paper no. 26.
  37. OECD (2013). OECD Skills Outlook 2013. First Results from the Survey of Adult Skills. OECD Publishing, Paris. http://skills.oecd.org/OECD_Skills_Outlook_2013.pdf
  38. OECD (2014). PISA 2012 Results in Focus. What 15-year-olds know and what they can do with what they know. Programme for International Student Assessment. Paris. http://www.oecd.org/pisa/keyfindings/pisa-2012-results.htm
  39. OECD and Statistics Canada (2000). Literacy in the information age. Final report of the International Adult Literacy Survey. Paris. http://www.oecd.org/edu/skills-beyond-school/41529765.pdf
  40. OECD and Statistics Canada (2011). Literacy for life: Further results from the Adult Literacy and Life Skills Survey. Second International ALL Report. OECD Publishing. http://www.statcan.gc.ca/pub/89-604-x/89-604-x2011001-eng.pdf
  41. OECD and Unesco Institute for Statistics (2003). Literacy Skills for the World of Tomorrow. Further results from PISA 2000. Programme for International Student Assessment. Paris. http://www.keepeek.com/Digital-Asset-Management/oecd/education/literacy-skills-for-the-world-of-tomorrow_9789264102873-en#page286
  42. Psacharopoulos, G. (1994), “Returns to investment in education: A global update.” World Development 22(9): 1325–1343.
  43. Schoellman, T. (2012). “Education quality and development accounting.” The Review of Economic Studies, 79(1), 388-417.
  44. Schwerdt, G. y S. Wiederhold (2018). “Literacy and Growth: New Evidence from PIAAC.” Mimeo, University of Konstanz and Catholic University Eichstaett-Ingolstadt
  45. Schwerdt, G. y S. Wiederhold (2019). “A Macroeconomic Analysis of Literacy and Economic Performance.” Mimeo.
  46. Wößmann, L. (2003). "Specifying human capital." Journal of Economic Surveys, 17(3), 239-270.