Nobel Prize in Economics 2000: Microeconometrics of Selective Samples and Discrete Choice
On this page
This note covers the Nobel Prize in Economics 2000: who won it, what microeconometrics is, how James Heckman solved the problem of selective samples, how Daniel McFadden built the theory of discrete choice, how their work developed from the 1970s onward, why it matters for policy and research today, and quick facts for exams.
What was the Nobel Prize in Economics 2000 awarded for?
The official citation, issued jointly for both laureates, reads: "for his development of theory and methods for analyzing selective samples" (James Heckman) and "for his development of theory and methods for analyzing discrete choice" (Daniel McFadden).
In plain words, both winners solved separate but related puzzles in microeconometrics, the branch of economics that uses statistics to study the behaviour of individual people, households or firms rather than whole economies.
Heckman worked out how to correct for the fact that the data we observe about people (such as wages) often come from a selective sample: we only see wages for people who chose to work, not for everyone.
McFadden worked out how to model choices that are not continuous numbers but a pick from a limited list of options, such as choosing to travel by car, bus or subway.
Together their tools let researchers use real-world data on individuals to test economic theories properly.
The prize's official name is the Sveriges Riksbank Prize in Economic Sciences in Memory of Alfred Nobel, commonly called the Nobel Prize in Economics. It was announced on 11 October 2000 and carried a prize amount of 9,000,000 Swedish kronor, split equally between the two laureates.
Who are the laureates?
James J. Heckman
James J. Heckman was born on 19 April 1944 in Chicago, IL, USA. At the time of the award he was affiliated with the University of Chicago, Chicago, IL, USA, where he held the Henry Schultz Distinguished Service Professorship of Economics from 1995. He received one half of the prize.
Heckman majored in mathematics at Colorado College before studying economics at Princeton University, where he earned his Ph.D. in 1971. He later taught at Columbia University and Yale University before joining Chicago.
His key contribution was developing statistical methods, most famously the Heckman correction (also called the two-stage method or Heckman's lambda), for handling data drawn from selective or self-selected samples, with major applications in labour economics, policy evaluation and duration analysis.
Daniel L. McFadden
Daniel L. McFadden was born on 29 July 1937 in Raleigh, NC, USA. At the time of the award he was affiliated with the University of California, Berkeley, CA, USA, where he held the E. Morris Cox Chair in Economics from 1990. He received one half of the prize.
McFadden studied physics as an undergraduate at the University of Minnesota, then moved into economics, earning his Ph.D. there in 1962.
He later taught at the University of Pittsburgh, Yale University and MIT before returning to Berkeley.
His landmark contribution was the conditional logit model of 1974, which gave discrete choice analysis a proper foundation in economic theory, with applications ranging from transport planning to environmental valuation.
What problem were they solving?
Before the 1970s, economists mostly tested theories using aggregate data, such as national income or total employment. From the late 1960s, large databases of microdata, meaning detailed information on individual people, households or firms, became available.
An early example is the Panel Study of Income Dynamics (PSID), which James Morgan and colleagues established at the University of Michigan in the late 1960s.
Growing computer power made it possible to analyse such data in new ways.
This opened up questions that aggregate statistics could never answer: what makes one person decide to work and for how many hours? How do education, training programmes or prices affect an individual's choices? But microdata brought new statistical headaches.
Two stood out. First, the data we can observe about people are often not a random sample of the population: we only see wages for those who chose to work, so if unobserved traits affect both wages and the decision to work, straightforward statistical analysis gives biased results.
Second, many important choices (occupation, travel mode, place of residence) are discrete, picked from a limited list of options, and classical demand theory, built for continuous quantities, could not handle this kind of choice.
Heckman tackled the first problem and McFadden the second, and both grounded their statistical fixes firmly in economic theory of how individuals maximise their own well-being.
How did Heckman solve the selective-sample problem?
A sample is called selective when the people we can observe are not a random slice of the whole population, often because the data depend on the individuals' own choices (self-selection).
For example, we can only observe a person's wage if that person chose to work.
If people with unusually high unobserved wage potential are more likely to work, then the sample of working people overstates how wages relate to education, understating the true effect.
The Nobel Prize's "Information for the Public" page explains that Heckman's breakthrough, developed in subsequent work in the mid to late 1970s building on his earlier studies of married women's labour supply, became known as the Heckman correction (also called the two-stage method or the Heckit method). The procedure can be summarised as:
- Build an economic model of the probability that an individual works, based on observed traits such as age and education.
- Estimate that participation model statistically to predict, for each individual, the chance that they work.
- Convert these predicted probabilities into a correction term (using the properties of the statistical distribution assumed for the unobserved factors).
- Add this correction term as an extra explanatory variable in the wage equation, then estimate the wage relationship in a way that is no longer distorted by self-selection.
Heckman also extended similar ideas to duration models, used to study how the chance of leaving unemployment changes the longer someone stays unemployed, and to the evaluation of labour-market programmes, where the challenge is to work out what would have happened to a participant had they not joined the programme.
The Nobel Prize's "Information for the Public" page notes that Heckman's own empirical findings on such programmes were "often quite pessimistic", since many showed only small positive, and sometimes negative, effects for participants.
Draw and label
wage against education with selection bias
Draw a scatter of points showing wage (vertical axis) against years of education (horizontal axis) for a whole population, with a solid line showing the true underlying relationship.
Then shade only the points for people whose wage lies above an assumed reservation wage (the dark, working subsample), and draw a second, flatter dashed line through just those points to show how the sample's visible relationship understates the true effect of education on wages.
How did McFadden build the theory of discrete choice?
McFadden's starting point was that many important decisions, such as which mode of transport to take to work, are choices among a small set of distinct options, not points on a continuous scale.
Standard demand theory, built for continuous quantities, could not be applied directly. His central 1974 paper, "Conditional Logit Analysis of Qualitative Choice Behavior", modelled each individual as choosing whichever available option gives them the highest utility, where utility depends partly on observed characteristics of the options and the individual, and partly on factors unobservable to the researcher.
By assuming that these unobserved factors follow a particular statistical distribution (an extreme-value distribution), McFadden showed that the probability of choosing a given alternative takes a clean mathematical form, now called the conditional logit model or multinomial logit model.
The committee noted this was "entirely new" compared with earlier, purely statistical logit models, because it was derived from genuine economic choice theory.
| Concept | What it describes |
|---|---|
| Conditional logit model | Predicts the probability an individual picks a given alternative from a fixed list, based on observed traits of the alternatives and the individual |
| Independence of irrelevant alternatives (IIA) | A restrictive property of the basic logit model: the relative odds of choosing between two options do not depend on any other option available |
| Nested logit model | Relaxes IIA by assuming choices are made in a sequence, such as choosing a location first and then a type of housing |
| Method of simulated moments | A simulation-based technique McFadden developed for estimating more general discrete-choice models when exact calculation is too complex |
McFadden recognised the basic logit model's weakness, the IIA property, meaning the model could behave oddly when a new, very similar alternative is added to the choice set.
He therefore devised statistical tests for IIA and developed more flexible alternatives, including the nested logit model (developed with Ben-Akiva) and his own generalized extreme value model, as well as simulation-based methods for even more general cases.
Draw and label
a simple nested choice tree
Draw a tree with a top box labelled "choice of location", branching into two lower boxes such as "City A" and "City B", and under each location box draw further branches for "type of housing", to show how a nested logit model breaks one big discrete choice into a sequence of smaller ones.
How did the discovery unfold?
| Year | Event |
|---|---|
| 1962 | Daniel McFadden earns his Ph.D. in economics from the University of Minnesota after switching from physics. |
| 1971 | James Heckman earns his Ph.D. in economics from Princeton University. |
| 1974 | Heckman publishes his study of married women's labour supply, introducing an econometric method for self-selection; McFadden publishes "Conditional Logit Analysis of Qualitative Choice Behavior", launching the conditional logit model. |
| 1976 to 1979 | Heckman develops the simpler, widely used two-stage correction method, later known as the Heckman correction or Heckit. |
| 1978 to 1981 | McFadden develops the nested logit and generalized extreme value models to relax the independence of irrelevant alternatives property. |
| 1984 | Heckman and Burton Singer propose a non-parametric estimator for duration models with unobserved heterogeneity; McFadden and co-authors apply discrete choice methods to residential energy demand. |
| 1989 | McFadden's discrete choice methods are later used to help assess environmental damage from the Exxon Valdez oil spill off Alaska. |
| 1995 | Heckman becomes Henry Schultz Distinguished Service Professor of Economics at the University of Chicago. |
| 2000 | Heckman and McFadden are jointly awarded the Sveriges Riksbank Prize in Economic Sciences in Memory of Alfred Nobel, announced on 11 October. |
Why does it matter?
The press release states that the tools developed by Heckman and McFadden "are now standard tools, not only among economists but also among other social scientists".
Before their work, researchers lacked reliable ways to use real individual-level data to answer practical questions: what determines whether someone works and for how many hours, how education and training programmes affect income, and what shapes choices such as travel mode or place of residence.
McFadden's own methods were applied to the design of the San Francisco BART (Bay Area Rapid Transit) system, as well as to studies of investment in telephone service and housing for the elderly, and to estimating environmental damage from the Exxon Valdez oil spill.
Heckman's correction became a routine tool in labour economics and in the evaluation of job-training and welfare programmes, and his duration-model methods are used throughout demography and labour economics to study unemployment spells, migration and fertility.
An open and important lesson from Heckman's programme-evaluation work, as the scientific background notes, is that there is no single universally correct method for evaluating social programmes: results depend heavily on the specific issue and the underlying economic model of participation and outcomes.
How does this connect to what you study?
Students who study statistics or economics at school will recognise some of these ideas in simpler form, even though neither laureate's methods are themselves school topics.
The problem of a biased sample, where the data collected do not represent the whole group being studied, is a basic statistics idea that appears whenever a survey only reaches people who chose to respond. This is exactly the problem for wages: researchers can only observe pay for people who decided to work, not for the whole population, so any simple comparison of wages and education risks being misleading unless this selection is accounted for.
Similarly, choosing between a fixed list of options, such as which subject to study, which career to pursue or which mode of transport to take to school, is exactly the kind of discrete choice that McFadden's conditional logit model describes mathematically, by assuming that each person picks whichever available option gives them the greatest personal benefit, given the information and preferences the researcher cannot fully observe.
A student's own understanding of probability, sampling and decision-making therefore connects directly to the real research recognised by this prize: both laureates showed that solid economic theory and careful statistics are needed together to draw correct conclusions from real human behaviour, a lesson useful well beyond economics itself.
Quick facts for exams
The Sveriges Riksbank Prize in Economic Sciences in Memory of Alfred Nobel 2000, widely known as the Nobel Prize in Economics 2000, was announced on 11 October 2000 and awarded jointly to James J. Heckman and Daniel L. McFadden, each receiving one half of the 9,000,000 Swedish kronor prize. Heckman, born in Chicago in 1944 and based at the University of Chicago, was honoured for developing theory and methods for analysing selective samples, including the Heckman correction.
McFadden, born in Raleigh, North Carolina in 1937 and based at the University of California, Berkeley, was honoured for developing theory and methods for analysing discrete choice, including the conditional logit model.
Together their work founded modern microeconometrics, the statistical study of individual economic behaviour.
| Fact | Detail |
|---|---|
| Prize | Sveriges Riksbank Prize in Economic Sciences in Memory of Alfred Nobel 2000 |
| Date announced | 11 October 2000 |
| Laureates | James J. Heckman and Daniel L. McFadden |
| Country of birth (both) | USA (Heckman: Chicago, IL; McFadden: Raleigh, NC) |
| Affiliation at award | Heckman: University of Chicago; McFadden: University of California, Berkeley, both USA |
| Share | One half each |
| Citation | "for his development of theory and methods for analyzing selective samples" (Heckman); "for his development of theory and methods for analyzing discrete choice" (McFadden) |
| Prize amount | 9,000,000 Swedish kronor |
Note: Source. The prize facts in this note are from the Nobel Prize's official site, nobelprize.org.
Glossary
- Microeconometrics — the branch of economics that uses statistical methods to analyse data on individuals, households or firms, rather than whole economies.
- Microdata — detailed data describing individual people, households or firms, as opposed to aggregate national statistics.
- Selective sample — a sample of data that does not randomly represent the underlying population, often because only certain individuals' information is observable.
- Self-selection — when the individuals included in a sample are there because of their own choices, such as choosing to work or to migrate.
- Selection bias — the distortion in statistical estimates that results from analysing a non-random, selective sample without correcting for it.
- Heckman correction — a two-stage statistical method developed by James Heckman to correct for self-selection bias, also called the Heckit method.
- Reservation wage — the minimum wage at which an individual is willing to work, according to standard labour-supply theory.
- Duration model — a statistical model used to study how long a state, such as unemployment, lasts and what affects the chance of it ending.
- Discrete choice — a choice made among a finite, limited set of distinct alternatives, such as travel mode or occupation.
- Conditional logit model — the statistical model developed by Daniel McFadden that predicts the probability of choosing each option in a discrete choice, grounded in utility-maximising economic theory.
- Independence of irrelevant alternatives (IIA) — a restrictive property of the basic logit model, under which the relative odds of choosing between two options are unaffected by other available options.
- Nested logit model — a more flexible discrete-choice model that organises choices into a sequence of stages to relax the IIA assumption.
- Method of simulated moments — a simulation technique McFadden developed for estimating complex discrete-choice models that are hard to calculate exactly.
Common errors and misconceptions
- Misconception: Heckman and McFadden worked on the same topic. Correct: Heckman's citation concerns selective samples, while McFadden's concerns discrete choice; they are related but distinct problems in microeconometrics.
- Misconception: The Heckman correction removes the need for good data. Correct: It corrects for a specific statistical bias caused by self-selection but still relies on a sensible economic model of participation.
- Misconception: Discrete choice means the choice itself is random. Correct: In McFadden's framework the individual's choice is treated as a deliberate, utility-maximising decision; the randomness is in factors the researcher cannot observe.
- Misconception: The conditional logit model always gives realistic predictions. Correct: The basic model has the restrictive independence of irrelevant alternatives property, which McFadden himself showed can be unrealistic, prompting the nested logit and other extensions.
- Misconception: This prize is only relevant to economists. Correct: The press release states the methods "are now standard tools, not only among economists but also among other social scientists".
- Misconception: Heckman's research always showed labour-market programmes work well. Correct: The Information for the Public page says his results were "often quite pessimistic", with many programmes showing only small or sometimes negative effects.
Exam-style questions with model answers
Q1. In which year was the Sveriges Riksbank Prize in Economic Sciences awarded to Heckman and McFadden? [1 mark]
- It was announced on 11 October 2000.
Q2. State the official citation for James Heckman's share of the prize. [2 marks]
- The citation reads: "for his development of theory and methods for analyzing selective samples".
Q3. What is meant by a "selective sample" in microeconometrics? [3 marks]
- A selective sample is one where the observed individuals do not randomly represent the whole population being studied. This often happens because of self-selection, where people choose to be included, for example only working individuals have observable wages. If the researcher ignores this, the estimated relationship between variables, such as education and wages, will be biased.
Q4. Briefly describe the Heckman correction and why it was needed. [4 marks]
- The Heckman correction is a two-stage statistical method for handling selection bias. It is needed because data such as wages are only observed for people who chose to work, so a direct estimate of the wage-education relationship using this subsample would be biased if unobserved factors affect both wages and the decision to work. In the first stage, the researcher models the probability of working and uses it to predict each individual's chance of working. In the second stage, this predicted probability is added as an extra variable in the wage equation, allowing the relationship to be estimated without the distortion caused by self-selection.
Q5. Explain McFadden's conditional logit model and give one real application mentioned in the sources. [4 marks]
- McFadden's conditional logit model treats an individual's choice among a limited set of alternatives as the outcome of utility maximisation, where utility depends on observed characteristics of the alternatives and the individual plus an unobserved random component. Assuming this random component follows an extreme-value distribution, McFadden showed the probability of choosing any alternative takes a convenient mathematical form. One real application is the design of the San Francisco BART transit system, where such models helped predict how people would choose between travel modes.
Q6. Discuss why microdata became important for economics from the late 1960s onward, and what new statistical problems this created. [5 marks]
- From the late 1960s, large databases recording detailed information on individuals and households, such as the Panel Study of Income Dynamics set up at the University of Michigan, became available alongside increasingly powerful computers. This allowed researchers to move beyond aggregate national statistics and test economic theories using data on real individual behaviour, answering questions such as what determines hours worked, how incentives affect choices of education or occupation, and what effects labour-market or educational programmes have. However, this microdata also created new statistical problems. Much of it reflects selective samples, since data such as wages are only observed for people who chose to work, and choices like occupation or travel mode are discrete rather than continuous, which traditional demand theory could not model properly. These two problems, selection bias and discrete choice, became the central challenges that Heckman and McFadden each resolved with methods now considered standard across economics and other social sciences.
Q7. Outline, in sequence, how the Heckman correction is applied to estimate a wage equation. [5 marks]
- First, the researcher specifies an economic model, based on theory, describing the probability that an individual chooses to work given observed traits such as age and education. Second, this participation model is estimated statistically across the whole sample, including both workers and non-workers, to generate a predicted probability of working for each individual. Third, this predicted probability is transformed into a correction term reflecting the likely influence of unobserved factors on the decision to work. Fourth, the correction term is added as an extra explanatory variable in the wage equation itself, alongside variables such as education. Fifth, the wage equation is then estimated using only the subsample of working individuals, but because the correction term accounts for the selection process, the resulting estimates of how education affects wages are no longer distorted by self-selection.
Q8. What criticism did McFadden identify in his own basic conditional logit model, and how did he address it? [3 marks]
- McFadden identified that the basic logit model has the independence of irrelevant alternatives property, meaning the relative odds of choosing between any two options stay fixed regardless of other options available, which is unrealistic when a new close substitute is introduced. He addressed this by developing more flexible models, including the nested logit model and the generalized extreme value model, which allow statistical dependence between related choices.
Key takeaways
- The 2000 Economics prize went jointly to James Heckman and Daniel McFadden for separate contributions to microeconometrics.
- Heckman's citation covers theory and methods for analysing selective samples, including his well-known two-stage correction.
- McFadden's citation covers theory and methods for analysing discrete choice, including the conditional logit model from 1974.
- Both laureates grounded their statistical innovations firmly in economic theory of individual decision-making.
- Heckman's methods are widely used in labour economics, policy evaluation and duration analysis of unemployment and demographic events.
- McFadden's methods underpin practical applications such as transport planning for the San Francisco BART system and environmental damage valuation.
- The prize carried 9,000,000 Swedish kronor, split equally, and was announced on 11 October 2000.
- The committee stated that both laureates' tools "are now standard tools" across economics and other social sciences.
Test yourself
Who shared the Nobel Prize in Economics 2000?
James J. Heckman and Daniel L. McFadden shared the prize equally, each receiving one half for separate contributions to microeconometrics.
Where was James Heckman based at the time of the award?
James Heckman was affiliated with the University of Chicago, Chicago, IL, USA, at the time of the award.
Where was Daniel McFadden based at the time of the award?
Daniel McFadden was affiliated with the University of California, Berkeley, CA, USA, at the time of the award.
What does "selective sample" mean in Heckman's work?
A selective sample means the data observed do not randomly represent the whole population, often because only people who chose to work, migrate or study are included.
What is the common name for Heckman's two-stage statistical fix?
It is widely known as the Heckman correction, also called the two-stage method or the Heckit method.
What kind of choices did McFadden's conditional logit model analyse?
It analysed discrete choices, meaning decisions made among a limited set of distinct alternatives, such as travel mode or occupation.
Name one real-world system whose design used McFadden's methods.
McFadden's methods were applied to the design of the San Francisco BART (Bay Area Rapid Transit) system.
What property of the basic logit model did McFadden try to relax with nested logit models?
He relaxed the independence of irrelevant alternatives property, which unrealistically keeps the relative odds between two choices unaffected by other options.
