Pseudogenes: Evolutionary Relics Hidden in DNA

The human genome contains thousands of genes that make proteins, regulate cellular processes, or help cells respond to their surroundings. It also contains DNA sequences that look remarkably like genes but no longer perform the same job. These sequences are called pseudogenes.

The name can be misleading. “Pseudo” means false, but pseudogenes are not fake genes or useless stretches of DNA. They are genetic sequences that resemble functional genes because they originated from them or from related genetic material, yet they generally lack the ability to produce a normal, functional gene product.

Some pseudogenes are evolutionary leftovers. Others have acquired new biological roles. Understanding them gives scientists a way to reconstruct the history of genomes while also revealing how DNA can be repurposed over evolutionary time.

What is a pseudogene?

A pseudogene is a DNA sequence that is recognizably related to a functional gene but has accumulated changes that prevent it from functioning like its ancestral or counterpart gene.

Those changes can include mutations that introduce premature stop signals, deletions or insertions that disrupt the reading frame, alterations that interfere with RNA production, or other changes that prevent the sequence from producing a functional protein.

A useful distinction is that a pseudogene is defined largely by its relationship to a functional gene or gene family, not simply by whether it currently produces a protein.

Pseudogenes can therefore look surprisingly similar to working genes. Some retain much of the original DNA sequence, including recognizable portions of coding regions. Others have become so altered that their connection to the original gene is detectable only through detailed genomic analysis.

How do pseudogenes form?

There are several major routes by which pseudogenes arise. The two most important are gene duplication and retrotransposition.

Duplicated genes can become pseudogenes

DNA duplication is a normal source of evolutionary innovation. Sometimes an organism acquires an extra copy of a gene. Because the original copy continues to perform the essential function, the duplicate can accumulate mutations without necessarily harming the organism.

Over generations, the duplicate may remain functional, develop a somewhat different function, or lose its ability to function as a gene. In the last case, it can become a duplicated pseudogene, sometimes called an unprocessed or non-retrotransposed pseudogene.

This process illustrates an important principle of evolution: genetic redundancy creates room for experimentation. A duplicated gene can tolerate changes that would be harmful if they occurred in the only working copy.

Not every duplicated pseudogene is completely inactive, however. A sequence that has lost its original protein-coding role may acquire regulatory activity or another function.

RNA can be copied back into DNA

Another route involves retrotransposition, a process in which an RNA molecule is converted back into DNA and inserted elsewhere in the genome.

When a messenger RNA is copied into DNA, it typically lacks the introns that were present in the original gene. Introns are sections of a gene that are removed from RNA during normal processing. The resulting DNA copy therefore has a characteristic structure: it resembles a gene but generally lacks the original gene’s introns and regulatory elements.

These sequences are called processed pseudogenes.

A processed pseudogene can also carry other clues to its origin. For example, because it originated from RNA, it may retain features associated with the original transcript rather than the original genomic gene.

Retrotransposition can produce many copies of related sequences throughout a genome. Most do not become functional genes, but occasionally a copied sequence acquires mutations that give it a useful new role.

Why do pseudogenes resemble real genes?

Their resemblance is a direct consequence of their history.

A pseudogene may have descended from a once-functional gene, so its DNA initially looked much like that gene. Evolution then modified the sequence through mutation, deletion, insertion, recombination, and other processes.

The degree of resemblance varies. A recently formed pseudogene may be almost indistinguishable from its functional relative except for a few disabling mutations. An ancient pseudogene may have accumulated so many changes that only sophisticated sequence comparisons reveal its ancestry.

Scientists can use these similarities to infer relationships among genes and reconstruct aspects of genome evolution.

Are pseudogenes actually useless?

No. The idea that every pseudogene is simply “junk DNA” is outdated.

Some pseudogenes appear to have little or no current biological function and can be considered evolutionary remnants. But other pseudogenes can be transcribed into RNA or otherwise influence gene regulation.

A pseudogene-derived RNA can sometimes interact with regulatory molecules or compete with other RNA molecules for binding partners. A pseudogene sequence can also influence the activity of a related functional gene through regulatory mechanisms.

In some cases, a pseudogene can contribute to biological function even though it does not encode a conventional protein.

This distinction matters because being unable to produce the original protein does not necessarily mean being biologically inactive.

At the same time, scientists should not assume that transcription automatically proves a function. Cells transcribe many genomic regions, and detecting an RNA molecule is not by itself evidence that the molecule has an important biological role.

Pseudogenes and the evolution of the human genome

Pseudogenes preserve traces of events that occurred during genome evolution.

Consider a gene that was duplicated in an ancestral population. If one copy remained essential while the other accumulated disabling mutations, the damaged copy could persist in the genome for millions of years. Its sequence would effectively record the existence of the ancestral duplication.

Comparing such sequences across species can reveal when particular genetic events occurred and how gene families changed over time.

Pseudogenes can therefore function as molecular fossils. They are not fossils in the geological sense, but they preserve evidence of earlier genetic states inside modern genomes.

They can also reveal the activity of genomic mechanisms such as retrotransposition and gene duplication. In this way, pseudogenes are evidence not only of what genes once did, but of how genomes continually copy, modify, rearrange, and repurpose their DNA.

How scientists identify a pseudogene

Finding a pseudogene is more complicated than looking for a broken gene.

Researchers typically compare DNA sequences with known or predicted genes and examine their structure and evolutionary relationships. A candidate pseudogene may share strong sequence similarity with a functional gene while containing features inconsistent with normal gene function.

Scientists may look for disrupted protein-coding regions, premature stop codons, frameshifts, missing regulatory sequences, unusual gene structures, or other evidence of gene inactivation.

Comparisons among species can strengthen the case. If a sequence is clearly related to a functional gene in other organisms but appears disabled in one lineage, that pattern can reveal when and how the loss of function occurred.

Functional experiments can then ask a different question: Does the pseudogene still do something?

That distinction is crucial. Identifying a sequence as a pseudogene describes its relationship to a functional gene and its loss of the original coding function. Determining whether the sequence has acquired another role requires additional evidence.

Why some pseudogenes are important in disease research

Pseudogenes have attracted attention in biomedical research because some are associated with changes in gene regulation and cellular behavior.

Cancer is one area in which researchers have investigated pseudogene-derived RNAs and their relationships with gene-expression networks. Some pseudogenes have also been studied in connection with other diseases and biological processes.

But an association does not necessarily mean that a pseudogene causes a disease. A pseudogene may influence a cellular process, respond to the same regulatory signals as another gene, or simply change its activity as a consequence of disease.

The broader lesson is that genome annotation and disease biology cannot always be reduced to a simple division between “functional genes” and “nonfunctional DNA.” Some genomic sequences can have regulatory effects that are very different from the protein-coding functions of their ancestors.

Pseudogenes are not all the same

The term pseudogene covers several evolutionary histories.

TypeHow it generally arisesTypical clue
Duplicated pseudogeneA DNA copy of a gene becomes nonfunctionalOften retains exon-intron structure
Processed pseudogeneAn RNA transcript is copied back into DNAUsually lacks the original introns
Unitary pseudogeneA once-functional gene loses its function without a surviving functional duplicateFunctional counterpart may be absent in that lineage

A unitary pseudogene is particularly informative about evolutionary history. In this situation, the organism has lost the function of a gene itself rather than simply accumulating a defective copy alongside a working one.

The consequences of such a loss depend on whether the gene’s function has become unnecessary, is compensated for by another biological mechanism, or was important only under particular environmental conditions.

What pseudogenes reveal about natural selection

Pseudogenes also provide a window into natural selection.

When a gene performs an important function, harmful mutations in that gene are often removed from a population by natural selection. Once a duplicate gene loses its essential function, however, mutations in that sequence may face much less selective pressure.

As a result, a nonfunctional copy can accumulate changes relatively freely.

This does not mean pseudogenes evolve in exactly the same way or at a perfectly constant rate. Some remain subject to selection because they retain a biological role, and genomic regions can experience different mutation and selection patterns. But the contrast between constrained functional genes and less-constrained pseudogene sequences can help researchers study evolutionary forces acting on DNA.

Can a pseudogene become a working gene again?

In principle, a pseudogene can acquire a new function, although this is not the same as simply repairing a broken gene.

Evolution can modify an existing sequence through additional mutations, regulatory changes, duplication, recombination, and other mechanisms. A formerly inactive sequence may therefore become biologically useful in a new way.

More commonly, a pseudogene-derived sequence can acquire a function different from that of its ancestor. In such cases, calling it an “evolutionary relic” captures only part of the story: the sequence may preserve an ancient history while participating in a newer biological process.

This is one reason genomic evolution is better understood as a process of continual modification than as a clean progression from useful DNA to useless DNA.

Why pseudogenes matter

Pseudogenes occupy an unusual position in the genome. They are often remnants of once-functional genetic sequences, yet their persistence can reveal how genomes change over time. Some are little more than molecular records of evolutionary events; others have been incorporated into regulatory systems or acquired functions of their own.

Their existence also illustrates an important feature of evolution: DNA does not have to be created from scratch for something new to evolve. Existing sequences can be copied, damaged, rearranged, silenced, and eventually repurposed.

For researchers, pseudogenes are therefore more than broken genes. They are clues to the history of gene families, evidence of genome-changing mechanisms, and reminders that the boundary between functional and nonfunctional DNA is more complicated than it first appears.

Looking For Something Else?