Human Genome Project

4 min read

On this page

Watch & explore

Start with a few high-quality watches, then dive into the notes below.

How to sequence the human genome - Mark J. Kiel · TED-Ed
The Human Genome Project Was a Failure · SciShow
The race to sequence the human genome - Tien Nguyen · TED-Ed

Try an idea before you read. Test your understanding of the Human Genome Project with this two-part scenario. Explore →

Why was the Human Genome Project started?

The Human Genome Project (HGP) was an international scientific effort that set out to read the base pairs that make up human DNA and to map every gene, working out what each part does and how it links to our physical features. It was the largest collaborative biological project in history. Planning began around 1990, and after roughly thirteen years of work the project was declared complete on 14 April 2003.

It was mainly a partnership between the United States Department of Energy and the National Institutes of Health. In its early years the Wellcome Trust in the United Kingdom was a major partner, and further funding and laboratory work came from Japan, France, Germany, China and others. In other words, dozens of countries pooled their scientists and money to read the human instruction manual for the first time.

Objectives of the Human Genome Project

The project had a clear set of goals:

  1. Identify all of the genes in human DNA, thought at the time to number between 20,000 and 25,000.
  2. Determine the sequence of the roughly 3 billion chemical base pairs that make up human DNA.
  3. Store this information in databases and build data analysis tools so that researchers everywhere could use it.
  4. Transfer the technologies it developed to the private sector so that they could be put to wider use.
  5. Study the ethical, legal and social issues (often shortened to ELSI) that such powerful knowledge might raise.

How can DNA sequencing help?

Learning the exact order of the letters in an organism's DNA gives us clues about what that organism can do. This knowledge can be applied to health care, agriculture, energy generation and cleaning up the environment, as well as helping us understand human biology itself. Once you can read the code, you can start to work out why some people are prone to certain diseases and how to treat them.

How was DNA sequenced for the Human Genome Project?

The project used two broad approaches to reading the genome:

  1. Expressed Sequence Tags (ESTs): this method focused only on the genes that are actually switched on and expressed as RNA, tagging and identifying them in a coordinated way.
  2. Random sequencing with sequence annotation: here the whole genome was sequenced at random, including both the coding and the non-coding stretches, and functions were then assigned to the different segments. Working out the job of each stretch of sequence is called sequence annotation.

How does DNA sequencing work?

DNA is an extremely long polymer, so reading very long sections at once is difficult. To get around this, the full DNA of a cell is broken into random, smaller fragments. Each fragment is then copied, or cloned, using special carriers called vectors.

Cloning multiplies each piece of DNA many times over, which makes it much easier and faster to sequence. The most common hosts were bacteria and yeast, and the vectors were known as BAC (bacterial artificial chromosomes) and YAC (yeast artificial chromosomes).

The fragments were then read automatically by DNA sequencers that followed the method developed by Frederick Sanger. Sanger is also credited with creating a way to work out the amino acid sequences of proteins.

Because the fragments overlapped in places, powerful computer programmes could line them up in the correct order, rather like rebuilding a shredded document by matching the torn edges. The ordered sequences were then annotated and assigned to their chromosomes. The sequence of chromosome 1 was only finished in May 2006, making it the last of the 24 human chromosomes, that is the 22 autosomes plus X and Y, to be completed.

A further challenge was matching the genetic map to the physical map of the genome. This was done using information about polymorphisms in restriction enzyme recognition sites and about microsatellites, which are short repeating DNA sequences.

Key findings of the Human Genome Project

  • The human genome contains about 3,164.7 million (roughly 3.1 billion) nucleotide bases.
  • A typical human gene is about 3,000 bases long. The longest known human gene, dystrophin, is about 2.4 million bases long.
  • The total number of genes was estimated at around 30,000, far fewer than the earlier guesses of 80,000 to 140,000.
  • Almost all humans share the same bases: about 99.9 per cent of our nucleotide bases are identical from person to person.
  • More than half of the genes found so far have functions we do not yet understand.
  • Less than 2 per cent of the genome actually codes for proteins.
  • A large part of the genome is made of repetitive sequences, which repeat many times. They do not code directly for proteins, but they tell us a lot about the structure, behaviour and evolution of chromosomes.
  • Chromosome 1 carries the most genes (2,968), while the Y chromosome carries the fewest (231).

Future applications of the Human Genome Project

Using DNA sequences to generate new knowledge will keep driving research and deepen our understanding of living systems. This is a vast task that draws on the skills of tens of thousands of scientists across many disciplines, in both the public and private sectors.

In the past, researchers studied one gene, or a few genes, at a time. Now that whole genome sequences and high throughput technologies exist, they can study problems on a much larger and more systematic scale. They can look at every gene in a genome at once, or every transcript in a particular tissue, organ or tumour, and see how thousands of genes work together in networks to run the chemistry of life.

The Human Genome Project also opened the door to many later projects, including HGP Write, the HapMap Project, the 100,000 Genomes Project and India's own Genome India Project. If you want to see how big ideas like this connect to careers and the wider world, the Learnacy Hub is a good place to explore next.

Why it still matters

The Human Genome Project is not just a chapter in a textbook. It is still shaping science today, and two recent developments show why.

First, the 2003 genome was never truly complete. Because of technical limits, scientists had only read about 92 per cent of the genome, leaving hard to sequence gaps. On 31 March 2022 an international group called the Telomere to Telomere (T2T) consortium published the first complete, gapless human genome sequence in the journal Science, filling in the last pieces nearly two decades after the HGP was declared finished. So the genome you learn about as finished in 2003 was actually only completed, letter for letter, in 2022.

Second, the original genome was mostly based on a small number of people, which does not capture the huge diversity of humans around the world. India has responded with the Genome India Project. Launched in 2020, it sequenced 10,000 whole genomes from 83 communities, with researchers drawn from about 20 institutes across the country. The government announced its completion in January 2025, and the data has been archived in India's national life science repository so that scientists can study diseases and traits that are specific to Indian populations. This is a real, current example of the HGP's legacy: a country building its own reference genome to make future medicine fairer and more accurate.

For students, the lesson is simple. Reading the genome was only the beginning. The work of understanding it, completing it and making it represent everyone is happening right now. You can read more topics like this in our science and technology study notes, or browse all of our revision notes.

Sources

  1. National Human Genome Research Institute, Telomere to Telomere programme: https://www.genome.gov/about-genomics/telomere-to-telomere
  2. Genome India Project (official portal): https://genomeindia.in/
  3. The Scientist, GenomeIndia Project: Mapping the Country's Genetic Diversity: https://www.the-scientist.com/india-maps-genomic-diversity-with-nationwide-project-72730

Key takeaways

  • The Human Genome Project was an international effort from 1990-2003 to map and sequence all human DNA, involving dozens of countries.
  • Its main goals were to identify all human genes, determine the sequence of roughly 3 billion base pairs, store the data, transfer technology to the private sector, and study ethical, legal, and social issues.
  • The project used two main approaches: Expressed Sequence Tags (ESTs) for expressed genes, and random sequencing with sequence annotation for the entire genome.
  • The human genome contains about 3.1 billion bases, with roughly 30,000 genes—far fewer than the earlier estimate of 80,000-140,000.
  • Almost all humans share 99.9% of their nucleotide bases, and less than 2% of the genome codes for proteins.

Test yourself

When was the Human Genome Project declared complete?

14 April 2003

How many genes were estimated in the human genome, and how does this compare to earlier guesses?

About 30,000 genes, which was far fewer than the earlier guesses of 80,000 to 140,000.

What percentage of the human genome actually codes for proteins?

Less than 2 per cent.

Frequently asked questions

What was the main goal of the Human Genome Project?

The project aimed to read all base pairs in human DNA, map every gene, and determine what each part does and how it links to physical features.

Why did the Human Genome Project sequence both coding and non-coding DNA?

Sequencing non-coding regions alongside genes helped assign functions to different segments through sequence annotation, providing a fuller picture of the genome’s organization.

How did the Human Genome Project make its data useful to researchers worldwide?

It stored the sequenced information in accessible databases and built data analysis tools so researchers everywhere could use the data.

What role did cloning play in DNA sequencing for the Human Genome Project?

Cloning multiplied DNA fragments using vectors like BACs and YACs, making it faster and easier to sequence the large number of copies needed for accurate reconstruction.

Try it

Human Genome Project

Test your understanding of the Human Genome Project with this two-part scenario.

1Scientists were surprised to discover that the human genome contains far fewer genes than originally expected. What was the initial estimate, and what did the project actually find?

2To sequence the human genome, DNA was broken into fragments, cloned, and then reassembled. How did scientists put the fragments back in the correct order?