Exons and Introns: Understanding Gene Structure

Genes are often described as stretches of DNA that contain instructions for making proteins. That description is useful, but it leaves out an important part of how genes actually work. In many genes, the DNA sequence that ultimately contributes to a protein is interrupted by segments that are removed before the protein-making instructions are used.

These two kinds of segments are called exons and introns. Understanding the difference helps explain how a gene is organized, how RNA is processed, how cells produce different proteins from the same gene, and how certain genetic changes can disrupt gene function.

What are exons and introns?

In a typical eukaryotic gene, the DNA region that is transcribed into RNA can contain alternating exons and introns.

Exons are sequences that remain in the mature RNA after RNA processing. In a protein-coding gene, the portions of exons that make up the protein-coding sequence are translated into amino acids. However, not every part of an exon necessarily codes for protein. Some exon sequences belong to untranslated regions (UTRs) at the beginning or end of the mature messenger RNA.

Introns are sequences that are transcribed into an initial RNA molecule but are removed during a process called RNA splicing. They therefore do not normally appear in the mature messenger RNA (mRNA) that is translated into a protein.

A simplified gene might look like this:

DNA:
Exon 1 — Intron 1 — Exon 2 — Intron 2 — Exon 3

After transcription, the cell initially produces a longer RNA containing both types of sequence. Splicing removes the introns and joins the exons:

Mature RNA:
Exon 1 — Exon 2 — Exon 3

The resulting RNA can then be used by the cell as a template for protein production if the gene encodes a protein.

How exons and introns fit into gene expression

Gene expression is the process by which information in DNA is used to produce a functional product, such as a protein or certain types of functional RNA.

For a protein-coding gene, the process begins when RNA polymerase copies the relevant DNA sequence into RNA. This first RNA product is called pre-mRNA, or precursor messenger RNA.

At this stage, the RNA still contains both exons and introns. The cell then processes the pre-mRNA. Among other modifications, it removes introns and joins neighboring exons through splicing.

The resulting mature mRNA contains the exon-derived sequences that will be used for translation, including the protein-coding region. The mRNA can then leave the nucleus and be read by ribosomes in the cytoplasm.

This means that the DNA organization of a gene and the structure of its final RNA are not identical. A gene can contain a long stretch of DNA that is transcribed but does not remain in the mature mRNA.

What is RNA splicing?

RNA splicing is the process of removing introns from pre-mRNA and connecting the remaining exon sequences.

A large molecular machine called the spliceosome carries out most splicing in eukaryotic cells. It is composed of RNA-protein complexes called small nuclear ribonucleoproteins, or snRNPs, along with additional proteins.

Splicing depends on sequence signals within the pre-mRNA that identify where introns begin and end. These signals help the cellular machinery distinguish introns from exons.

The process is more than simply cutting out a piece of RNA. The intron is removed through a characteristic reaction in which the intron forms a temporary loop-like structure called a lariat. The two neighboring exons are then joined together.

Accurate splicing is essential because even a small error at an exon-intron boundary can alter the final RNA sequence and, consequently, the protein produced from it.

Exons do not always mean “protein-coding DNA”

One of the most common oversimplifications is the idea that exons are simply the parts of a gene that code for a protein.

That is not quite correct.

An exon is defined by its fate during RNA processing: it is retained in the mature RNA. In a protein-coding gene, exons can contain both coding and noncoding portions.

For example, mature mRNA generally contains a 5′ untranslated region (5′ UTR) before the protein-coding sequence and a 3′ untranslated region (3′ UTR) after it. These regions are not translated into protein, but they can play important roles in controlling mRNA stability, localization, and translation.

Thus, the terms have different meanings:

  • Exon: a sequence retained in the mature RNA.
  • Intron: a sequence removed from the pre-mRNA during splicing.
  • Coding sequence: the portion of an mRNA that is translated into a protein.

These categories overlap in some cases but are not interchangeable.

Why do genes contain introns?

Introns were once sometimes portrayed as useless stretches of DNA that interrupt useful sequences. Modern molecular biology provides a more nuanced picture.

Introns are not generally translated into protein, but they can participate in gene regulation and RNA processing. Their presence also enables extensive alternative splicing, allowing cells to generate different mature RNA molecules from the same gene.

Introns can contain regulatory elements and, in some cases, sequences that give rise to functional RNAs. Their removal can also influence aspects of gene expression beyond simply determining which nucleotides remain in the mature mRNA.

At the same time, not every intron has a known independent function. The biological roles and evolutionary histories of introns vary among genes and organisms.

Alternative splicing: one gene, multiple RNA products

One of the most important consequences of exon-intron organization is alternative splicing.

Instead of always joining every exon in exactly the same way, cells can sometimes splice a pre-mRNA differently. An exon may be included in one mature mRNA but skipped in another, or different splice sites may be selected.

For example:

Pre-mRNA:
Exon 1 — Exon 2 — Exon 3 — Exon 4

One cell type might produce:

mRNA A:
Exon 1 — Exon 2 — Exon 3 — Exon 4

Another might produce:

mRNA B:
Exon 1 — Exon 3 — Exon 4

If the resulting coding sequences remain in the appropriate reading frame, the two mRNAs can produce related proteins with different structures or functions.

Alternative splicing is especially important in multicellular organisms, where different tissues and developmental stages can require different versions of gene products.

The phrase “one gene, one protein” therefore does not accurately describe many eukaryotic genes. A single gene can give rise to multiple RNA and protein products through alternative processing, although the exact number and biological significance of these products vary by gene.

What happens when an exon or intron contains a mutation?

Changes in DNA can affect exons, introns, or the sequences that control splicing.

A mutation within a protein-coding portion of an exon may change the resulting protein. Depending on the particular change, it can substitute one amino acid for another, introduce a premature stop signal, or otherwise alter the protein-coding sequence.

Mutations affecting introns can also matter. An intron contains sequence information that helps define where splicing should occur. A mutation near an exon-intron boundary, for example, can interfere with normal recognition of a splice site.

The consequences can include exon skipping, retention of an intron, or use of an incorrect splice site. Any of these changes can alter the mature mRNA and potentially produce an abnormal or nonfunctional protein.

This is why it is misleading to assume that a mutation is harmless simply because it lies outside a protein-coding exon. Noncoding portions of genes can contain important information for gene expression and RNA processing.

Exon-intron boundaries are especially important

The boundaries between exons and introns contain signals that guide the splicing machinery.

A typical intron has characteristic sequence features near its ends, including a 5′ splice site and a 3′ splice site, as well as an important internal branch-point sequence. The exact sequences and mechanisms vary, and splicing is regulated by a broader collection of RNA and protein signals.

Because these regions help determine where an intron should be removed, mutations that disrupt them can have substantial effects on gene expression.

Genetic testing therefore sometimes examines splice-site variants alongside changes within protein-coding regions. Determining whether a particular variant actually alters splicing may require additional experimental or computational evidence.

Exons and introns in DNA versus RNA

The terminology can become confusing because exons and introns are often discussed as though they exist only in RNA. They are defined in relation to RNA processing, but their corresponding sequences are encoded in the gene’s DNA.

A DNA sequence that will function as an intron is copied into the pre-mRNA and subsequently removed. A DNA sequence corresponding to an exon is copied into the RNA and retained in the mature transcript.

This distinction is useful:

In the gene’s DNA: exon and intron regions are physically present as part of the gene.

In pre-mRNA: both exon and intron sequences are present.

In mature mRNA: introns have been removed, while exon-derived sequences remain.

The same underlying DNA can therefore give rise to RNA molecules with a different arrangement of sequences.

Are all genes made of exons and introns?

No.

The exon-intron arrangement described above is characteristic of many eukaryotic genes, including genes in humans and other animals, plants, fungi, and many other eukaryotes. But gene structures vary widely across biology.

Some genes contain no introns. Others contain many introns, and intron lengths can vary substantially. Some genes also produce functional RNAs rather than proteins, so their exons are not necessarily protein-coding.

Prokaryotes, such as bacteria, generally have much more compact genomes and typically lack the widespread spliceosomal intron structure found in eukaryotic protein-coding genes. There are exceptions and other forms of RNA processing, but the familiar exon-intron model is primarily associated with eukaryotic gene expression.

Why exon and intron structure matters

The arrangement of exons and introns gives cells another layer of control over genetic information. DNA is not simply copied into a final RNA molecule and immediately translated. Instead, the initial transcript can be processed in ways that determine which sequences appear in the mature RNA.

This organization helps explain several important features of molecular genetics: why a gene can produce multiple related proteins, why changes in noncoding DNA can sometimes cause disease, and why understanding a genetic variant often requires looking at RNA processing as well as the protein-coding sequence.

In short, exons are sequences retained in mature RNA, whereas introns are sequences removed during RNA splicing. For protein-coding genes, the retained exons provide the framework for the mature messenger RNA, including its coding and untranslated regions. The regulated removal and joining of these sequences is a central part of gene expression in eukaryotic cells.

Looking For Something Else?