RNA splicing is a crucial step in gene expression. It converts a newly made RNA transcript into a form that can be used to produce a protein or perform other cellular functions. The central task is simple to state: remove introns and join exons together.
The process is more precise than that shorthand suggests. Splicing must recognize the correct boundaries within an RNA molecule, cut at exactly the right positions, and connect the remaining pieces in the proper order. Most pre-mRNA splicing in human cells is carried out by a large molecular machine called the spliceosome.
Understanding splicing helps explain how a single gene can produce different RNA and protein products, why mutations in noncoding regions can cause disease, and how cells control which genetic instructions ultimately reach the protein-making machinery.
What is RNA splicing?
When a protein-coding gene is expressed, its DNA sequence is first copied into a preliminary RNA molecule called pre-messenger RNA (pre-mRNA). In many human genes, this initial transcript contains two kinds of sequence:
- Exons, which are the segments retained in the mature RNA
- Introns, which are removed during RNA processing
RNA splicing removes the introns and joins the exons. The resulting mature mRNA can then leave the nucleus and, in the case of protein-coding mRNA, be translated by ribosomes.
For example, imagine a pre-mRNA arranged like this:
Exon 1 — Intron 1 — Exon 2 — Intron 2 — Exon 3
After splicing, the mature RNA becomes:
Exon 1 — Exon 2 — Exon 3
The introns do not normally become part of the final protein-coding sequence.
Splicing occurs primarily in the nucleus, where pre-mRNA is produced. It is part of a broader RNA-processing pathway that also includes processes such as addition of a 5′ cap and formation of the 3′ end of the RNA.
Why do genes contain introns?
Introns are sometimes described as genetic “junk,” but that description is misleading. Introns are generally not translated into protein, but they can contain regulatory sequences and can influence how genes are expressed and spliced.
More importantly, the presence of introns gives cells additional ways to regulate gene expression. The same pre-mRNA can sometimes be spliced in different ways, allowing a single gene to produce multiple mature RNA molecules.
This process is called alternative splicing.
The distinction between introns and exons is also about RNA processing rather than simply function. An exon is defined as a sequence retained in a mature RNA product, while an intron is a sequence removed from that RNA during splicing. An exon does not necessarily consist entirely of protein-coding sequence; mature mRNAs also contain untranslated regions.
How does the spliceosome know where to cut?
The spliceosome does not remove introns randomly. It recognizes characteristic sequence features near intron boundaries.
A typical pre-mRNA intron contains three especially important elements:
- A 5′ splice site at the beginning of the intron
- A branch point within the intron
- A 3′ splice site near the end of the intron
The exact sequences vary, but these regions provide molecular signals that help the splicing machinery identify the boundaries.
The branch point contains an important adenosine nucleotide. Its position relative to the splice sites is important because this adenosine supplies the chemical group used to form a distinctive looped structure during splicing.
A region rich in pyrimidine nucleotides, called the polypyrimidine tract, is also commonly found between the branch point and the 3′ splice site in many pre-mRNAs.
These sequence signals do not operate alone. Proteins and small RNA-containing complexes work together to identify appropriate splice sites and assemble the splicing machinery.
The spliceosome: the cell’s splicing machinery
The spliceosome is a dynamic molecular complex made primarily of proteins and small nuclear RNAs, or snRNAs.
Several major spliceosomal components are known as small nuclear ribonucleoproteins, or snRNPs. The commonly discussed major spliceosome contains U1, U2, U4, U6, and U5 snRNPs.
The names refer to the small nuclear RNAs within these complexes. Each snRNP also contains associated proteins.
Rather than functioning as a rigid machine with fixed parts, the spliceosome assembles and rearranges on the pre-mRNA as splicing proceeds. RNA-RNA and RNA-protein interactions help identify splice sites, position the relevant nucleotides, and carry out the chemical reactions.
How an intron is removed
Splicing can be understood as a sequence of recognition, rearrangement, and chemical reactions.
1. The splice sites are recognized
The spliceosome first identifies important features of the intron and its boundaries. U1 snRNP interacts with the 5′ splice site, while U2 snRNP recognizes the branch-point region.
Other components are then recruited, producing a larger spliceosomal complex.
The spliceosome undergoes substantial structural rearrangement before the RNA-cutting reactions occur. This rearrangement is important because the machinery must place the reactive RNA groups in the correct geometry.
2. The intron forms a loop
The first major chemical reaction involves the branch-point adenosine.
Its 2′ hydroxyl group attacks the phosphate at the 5′ splice site. This breaks the bond connecting the intron to the upstream exon and simultaneously creates a new bond between the branch-point adenosine and the beginning of the intron.
The result is a characteristic looped structure called a lariat.
The intron is now attached to itself at the branch point, giving it a shape somewhat like a loop with a tail.
3. The exons are joined
In the second major reaction, the newly exposed end of the upstream exon attacks the phosphate at the 3′ splice site.
This joins the upstream and downstream exons together.
The intron lariat is released as a separate RNA molecule. It is subsequently dismantled and degraded, although some intron-derived RNAs can have additional functions.
The mature RNA now contains the correctly joined exons.
What actually drives the chemical reactions?
A useful point about splicing is that the two central reactions are transesterification reactions. They rearrange phosphodiester bonds in the RNA backbone.
Because one phosphodiester bond is exchanged for another during each reaction, the chemistry itself does not require the spliceosome to supply energy in the same way an enzyme-driven bond-breaking process might.
However, ATP-dependent molecular rearrangements are important for the spliceosome as a whole. Proteins in the spliceosome use ATP to remodel RNA-protein interactions and change the complex’s structure during the splicing cycle.
So ATP is essential to the operation and regulation of the splicing machinery even though the two core transesterification reactions themselves do not directly require ATP hydrolysis.
What happens to the intron afterward?
Once the exons have been joined, the intron is released in lariat form. The lariat is typically debranched, meaning the unusual branch-point linkage is broken, and the resulting RNA is degraded and its components recycled.
The cell therefore separates the intron from the mature mRNA rather than simply cutting it out as a straight piece.
This lariat intermediate is one of the defining molecular features of conventional spliceosomal intron removal.
Alternative splicing lets one gene produce different RNAs
Not every pre-mRNA is spliced in exactly the same way.
Alternative splicing allows cells to select different combinations of exons or use different splice sites. As a result, one gene can give rise to multiple mature RNA isoforms.
For instance, a pre-mRNA might contain three exons, but a particular cell type could produce an RNA containing all three, while another context produces an RNA in which one exon is skipped.
Alternative splicing can therefore change the protein produced from a gene. Depending on which exons are retained, the resulting protein may have different domains, altered regulatory properties, or different cellular destinations.
Splicing patterns can vary among tissues and developmental stages and can change in response to cellular conditions. This makes splicing an important layer of gene regulation rather than merely a housekeeping step that removes unwanted RNA.
Splicing is closely tied to gene expression
Splicing does not necessarily happen only after transcription is complete. In many cases, spliceosome assembly and intron removal occur while the RNA is still being transcribed.
This coordination is known as co-transcriptional splicing.
The connection between transcription and splicing means that the timing of RNA polymerase movement, the structure of the emerging RNA, and the availability of regulatory proteins can influence how splice sites are selected.
Proteins called splicing factors help regulate these choices. Some promote recognition of particular splice sites, while others discourage their use. The relative amounts and activities of these factors help determine which RNA isoforms a cell produces.
What happens when splicing goes wrong?
Accurate splicing is essential because even a small change in splice-site selection can alter the final RNA.
A mutation can disrupt a normal splice site, create a new potential splice site, or alter regulatory sequences that control splice-site selection. The result may be an mRNA missing an exon, retaining part of an intron, or using an inappropriate splice boundary.
These changes can affect the protein in several ways. They may alter its amino acid sequence, remove an important functional region, introduce a premature stop signal, or prevent production of a stable protein altogether.
Splicing abnormalities are therefore involved in many human diseases, including inherited disorders and cancers. Importantly, the mutation responsible for a splicing defect does not have to occur directly at the most obvious splice-site sequence. Changes in other regulatory regions can also disturb normal splice-site choice.
The key idea: splicing changes RNA, not DNA
RNA splicing does not remove introns from the chromosome itself.
The DNA remains unchanged. Instead, the cell first copies the relevant gene into an RNA transcript containing both exons and introns. The spliceosome then processes that RNA transcript by removing introns and joining exons.
This distinction matters because the same DNA sequence can continue to serve as the template for future RNA molecules.
Splicing is therefore an example of how cells can regulate genetic information after DNA has been transcribed, adding another layer of control between a gene and its final biological product.
