Genes are often described as stretches of DNA that contain instructions for making proteins. That description is useful, but incomplete. In many human genes, the protein-coding information is interrupted by segments that are removed before the cell uses the gene’s RNA instructions to make a protein.
These two kinds of segments are called exons and introns. Exons are sequences retained in the mature RNA, while introns are sequences removed from the RNA during a process called RNA splicing. Understanding how they fit together helps explain how DNA is converted into functional RNA and, for protein-coding genes, ultimately into proteins.
The distinction is also important because exons and introns do more than divide a gene into simple “useful” and “unused” pieces. Some exons contain protein-coding information, but others contribute to RNA regions that are not translated into protein. Introns can contain regulatory information and other functional sequences as well.
Where exons and introns fit in gene expression
For a typical protein-coding gene, information flows through several stages:
DNA → pre-mRNA → mature mRNA → protein
The gene is encoded in DNA. When the gene is expressed, the DNA sequence is copied into an initial RNA molecule called pre-mRNA, or precursor messenger RNA. This RNA generally contains both exons and introns.
Before the RNA can function as a mature messenger RNA, the cell processes it. One of the most important steps is splicing, during which introns are removed and exons are joined together.
The resulting mature mRNA can then be used by ribosomes as a template for protein production. The protein itself is not assembled directly from DNA or from the original pre-mRNA; the cell uses the processed RNA.
A simplified example looks like this:
DNA:
Exon 1 — Intron 1 — Exon 2 — Intron 2 — Exon 3
Pre-mRNA:
Exon 1 — Intron 1 — Exon 2 — Intron 2 — Exon 3
Mature mRNA:
Exon 1 — Exon 2 — Exon 3
The introns have been removed, and the remaining exons have been connected into a continuous RNA molecule.
What is an exon?
An exon is a portion of a gene that remains in the mature RNA after splicing.
In a protein-coding gene, some or all of an exon may contain information that is translated into a protein. But “exon” and “protein-coding sequence” are not synonymous.
For example, mature messenger RNA usually contains regions at both ends that are not translated into protein. These are called untranslated regions, or UTRs. They are part of the mature mRNA and therefore can be derived from exons, even though they do not specify amino acids.
This distinction matters because researchers use several different terms to describe different aspects of a gene:
- Exon: a sequence retained in the mature RNA.
- Coding sequence: the portion of an mRNA that is translated into amino acids.
- UTR: an untranslated portion of the mature mRNA that can have important regulatory roles.
Thus, an exon can contain coding sequence, UTR sequence, or both.
What is an intron?
An intron is a segment of a gene that is transcribed into the initial RNA but removed during RNA processing.
Introns therefore occupy a different position from DNA sequences that are never transcribed. An intron is initially copied into RNA; it is subsequently removed through splicing.
Introns vary considerably in size and sequence. They are not simply meaningless stretches of DNA. Some contain regulatory elements or sequences that can influence gene expression and RNA processing. Their presence also makes it possible for cells to process a single gene’s RNA in different ways.
Because introns are removed from mature messenger RNA, they generally do not appear in the final mRNA sequence between its exons.
How RNA splicing removes introns
RNA splicing is carried out by a molecular system called the spliceosome, along with associated proteins and RNA molecules.
The spliceosome recognizes characteristic sequence signals around introns. It then carries out a precisely coordinated series of reactions that removes the intron and joins the neighboring exons.
The boundaries between exons and introns are therefore biologically important. A mutation that disrupts a splice site can cause the cell to splice an RNA incorrectly. Depending on the gene and the particular mutation, the resulting mRNA may contain an unwanted sequence, lose part of an exon, or otherwise produce an abnormal transcript.
Splicing is consequently not just a cleanup step. It is a tightly regulated part of gene expression.
Alternative splicing allows one gene to produce different RNAs
Not every transcript from a gene has to contain exactly the same set of exons.
Through alternative splicing, cells can combine certain exons in different patterns. One mature RNA might contain exons 1, 2, and 4, for example, while another transcript from the same gene might contain exons 1, 3, and 4.
This can allow a single gene to produce multiple related RNA molecules and, in protein-coding genes, potentially different protein isoforms.
Alternative splicing is regulated rather than random. Proteins that bind RNA can influence whether particular splice sites are used, and splicing patterns can vary among cell types or biological conditions.
This is one reason that the relationship between genes and proteins is more complicated than a simple one-gene-to-one-protein model.
Exons do not always correspond to one complete protein feature
It is tempting to think of each exon as encoding one particular part of a protein, but that is not generally correct.
A protein-coding region can span several exons, with individual exons contributing portions of the final amino-acid sequence. An exon boundary can occur within a protein-coding sequence rather than neatly separating one functional part of a protein from another.
Likewise, some exons may include untranslated sequence at one or both ends of the coding region.
The biological significance of an exon therefore depends on its exact position and sequence, not simply on the fact that it is an exon.
What happens when DNA is changed at an exon or intron?
The effect of a genetic variant depends heavily on where it occurs and how it affects the gene.
A change within a protein-coding portion of an exon can alter the resulting protein. Depending on the particular change, it may substitute one amino acid for another, introduce a premature stop signal, or have little or no effect on the protein.
Variants in exons can also affect splicing. A sequence change may interfere with normal splice-site recognition or alter regulatory sequences that help control exon inclusion.
Intronic variants can have effects as well. A change near an intron-exon boundary may disrupt normal splicing, while changes farther into an intron can sometimes affect regulatory or other functional elements.
For this reason, describing a variant simply as “exonic” or “intronic” does not by itself determine whether it is biologically important.
Why gene structure matters in genetics and medicine
The exon-intron organization of genes is central to how genetic variants are interpreted.
When scientists analyze a DNA sequence, they need to know which portions correspond to exons, introns, regulatory regions, and other genomic features. A variant in a protein-coding exon may immediately raise questions about its effect on the amino-acid sequence, while a variant affecting a splice site raises a different question: does the mutation change how the RNA is assembled?
Some genetic disorders result from mutations that alter splicing rather than directly changing the protein-coding sequence. The mutation may cause an exon to be skipped, cause part of an intron to be retained, or activate an abnormal splice site. The resulting RNA can then encode an altered or nonfunctional protein.
This is why modern genetic analysis considers both the DNA sequence itself and the way that sequence is used to produce RNA.
Exons and introns in context
The simplest useful picture is that exons are retained and introns are removed during RNA splicing. But gene structure is more nuanced than that shorthand suggests.
A protein-coding gene can contain multiple exons separated by introns. The entire region can be transcribed into a precursor RNA, after which splicing and other RNA-processing steps produce a mature transcript. Exons can contain both coding and untranslated sequences, while introns can contain sequences with regulatory or other functions.
The distinction also applies differently across types of genes and transcripts. Not every gene produces a protein-coding mRNA, and RNA molecules can have structures and processing pathways that do not fit the simplest textbook diagram.
At its core, however, the concept is straightforward: introns are intervening sequences removed from a precursor RNA, while exons are sequences retained in the mature RNA. Their arrangement gives cells another layer of control over how genetic information is expressed, helping explain how the same genome can support the diverse cells and functions of the human body.

