Every cell must solve a fundamental problem: how does it turn information stored in DNA into the molecules that carry out most of its work?
The answer is a chain of information flow. In broad terms, cells store genetic information in DNA, copy selected DNA sequences into RNA, and use that RNA to direct the production of proteins. This is often summarized as:
DNA → RNA → protein
The pathway is more nuanced than that shorthand suggests. DNA is not simply read from beginning to end, and not every gene produces a protein. Cells regulate which genes are used, process many RNA molecules before they become functional, and use proteins alongside RNAs to control nearly every stage of the process.
Understanding this flow provides a foundation for understanding heredity, development, disease, biotechnology, and many of the basic processes of life.
DNA stores the instructions
DNA, or deoxyribonucleic acid, is the cell’s long-term repository of genetic information. In humans and other eukaryotes, most DNA is contained in chromosomes inside the nucleus, although mitochondria also contain their own DNA.
DNA consists of four chemical bases: adenine (A), thymine (T), cytosine (C), and guanine (G). The order of these bases carries information. Because A pairs with T and C pairs with G, the two strands of the DNA molecule can serve as complementary templates for one another.
A gene is a segment of DNA that contains information used to produce a functional biological product, usually a protein but sometimes a functional RNA molecule. Genes do not exist as isolated instructions, however. They are embedded in larger stretches of DNA that include regulatory sequences and other regions that help determine when, where, and how strongly a gene is used.
The sequence of bases in a gene ultimately determines the sequence of amino acids in a protein, but there is an important intermediate step: RNA.
Transcription copies genetic information into RNA
The first major step in using a protein-coding gene is transcription. During transcription, an enzyme called RNA polymerase uses one strand of DNA as a template to build an RNA molecule.
RNA resembles DNA but has several important differences. RNA is usually single-stranded, contains the sugar ribose rather than deoxyribose, and uses uracil (U) in place of thymine. When RNA polymerase reads a DNA template, it builds an RNA sequence according to base-pairing rules: DNA A corresponds to RNA U, while DNA T corresponds to RNA A; C pairs with G and G with C.
The initial RNA produced from a protein-coding gene in a eukaryotic cell is called a pre-mRNA, or precursor messenger RNA. It is not necessarily ready to be used by a ribosome.
Transcription is also highly regulated. A cell does not transcribe every gene at all times. Different cell types activate different sets of genes, and even the same cell can change gene activity in response to signals such as hormones, nutrients, stress, or changes in its environment.
This selective use of genes is one reason cells with essentially the same genome can have dramatically different structures and functions. A neuron and a muscle cell, for example, use different subsets of their genetic information.
Eukaryotic RNA is processed before it leaves the nucleus
In eukaryotic cells, which include human cells, newly transcribed pre-mRNA generally undergoes several processing steps before it becomes mature messenger RNA, or mRNA.
One modification is the addition of a 5′ cap to the beginning of the RNA molecule. Another is the addition of a poly(A) tail to its other end. These structures help with RNA stability, transport, and recognition by the machinery involved in protein production.
A particularly important step is RNA splicing. Protein-coding genes in eukaryotes commonly contain segments called exons, which remain in the mature RNA, and introns, which are removed. The cell’s RNA-processing machinery cuts out the introns and joins the exons together.
Splicing also allows some genes to produce different mRNA molecules by combining exons in different ways. This process, called alternative splicing, expands the range of proteins that can be produced from a genome.
Once properly processed, many mRNAs are transported from the nucleus into the cytoplasm. There, their nucleotide sequence can be interpreted by ribosomes.
Translation turns an RNA sequence into a protein
The second major information-transfer step is translation. During translation, a ribosome reads the sequence of an mRNA and uses it to assemble a chain of amino acids.
Proteins are made from 20 common amino acids. The ribosome does not read the mRNA one nucleotide at a time. Instead, it reads it in groups of three nucleotides called codons.
Because there are four possible RNA bases and three positions in each codon, there are 64 possible codons. Most specify amino acids, while some signal the end of translation. The genetic code is therefore degenerate, meaning that more than one codon can specify the same amino acid.
Transfer RNA, or tRNA, helps connect codons with their corresponding amino acids. Each tRNA carries a particular amino acid and contains an anticodon that can pair with a complementary codon in the mRNA.
As the ribosome moves along the mRNA, it links the appropriate amino acids together. The result is a polypeptide, a chain of amino acids whose sequence is determined by the mRNA.
The sequence of amino acids is crucial because it influences how the polypeptide folds and functions.
A protein’s amino acid sequence helps determine its shape
A newly synthesized polypeptide is not necessarily a finished, functional protein. It must generally fold into a particular three-dimensional structure, and some proteins undergo additional chemical modifications or are assembled with other protein molecules.
The amino acid sequence influences folding through chemical interactions among the amino acids and between the chain and its surroundings. A protein’s resulting shape allows it to interact selectively with other molecules.
This structure-function relationship is central to biology. Proteins can act as enzymes that accelerate chemical reactions, receptors that detect signals, channels that move substances across membranes, structural components of cells, molecular motors, antibodies, and regulators of gene activity.
A change in DNA can therefore have effects far beyond the DNA molecule itself. If a DNA variant changes an mRNA sequence, it may alter a codon and consequently change an amino acid in the resulting protein. Depending on the location and nature of the change, the protein’s function may be unaffected, altered, or lost.
Not every DNA change affects a protein. Some changes occur in regions that do not encode amino acids, some alter synonymous codons without changing the amino acid sequence, and some influence how much or when a gene is expressed rather than changing the protein’s sequence.
The information flow is regulated at many levels
The path from DNA to protein is not a simple assembly line operating at a fixed rate. Cells control gene expression at multiple stages.
Before transcription begins, regulatory proteins and DNA sequences can influence whether a gene is accessible to the transcription machinery. Chemical modifications to DNA and associated proteins can also affect how readily particular regions of the genome are used. These mechanisms are part of gene regulation and epigenetic regulation.
After transcription, the cell can regulate how an RNA molecule is processed, how long it survives, and whether it is transported or translated efficiently. Small regulatory RNAs can also influence the stability or translation of particular mRNAs.
Translation itself can be regulated, as can the activity, location, modification, and degradation of the resulting proteins.
This layered control allows cells to adjust protein production with considerable precision. A cell does not merely ask, “Is this gene present?” It must regulate questions such as when should this gene be active, in which cells, at what level, and for how long?
Not all genetic information follows the DNA → RNA → protein path
The phrase “central dogma” is often used to describe the general flow of biological information from DNA to RNA to protein. It should not be interpreted as meaning that all information transfer in a cell follows only that route.
Some genes produce functional RNA molecules rather than proteins. Ribosomal RNA (rRNA) and transfer RNA (tRNA) are examples, as are several other regulatory and structural RNAs.
There are also biological processes in which RNA serves as a template for producing DNA. Reverse transcription, for example, occurs in certain viruses and is also part of the life cycle of retroelements. These processes do not overturn the basic principle that the sequence information used to specify a protein is read through an RNA intermediate; they illustrate that biological information flow is more diverse than the simplest diagram suggests.
Where the process happens in a human cell
In a typical human cell, the major stages are divided between the nucleus and cytoplasm.
Transcription takes place in the nucleus, where the cell’s chromosomes are located. RNA processing also occurs primarily in the nucleus. Mature mRNA can then pass through nuclear pores into the cytoplasm.
Translation occurs on ribosomes in the cytoplasm. Some ribosomes are free in the cytosol, while others are attached to the rough endoplasmic reticulum. Where a protein is synthesized helps determine where it will ultimately function.
Proteins destined for secretion, for insertion into cellular membranes, or for certain compartments of the endomembrane system are commonly synthesized on ribosomes associated with the rough endoplasmic reticulum. Other proteins are produced on free ribosomes and can remain in the cytosol or be directed to other cellular locations.
Thus, genetic information is not simply converted into a protein and left there. Cells also have systems that direct proteins to the right place, modify them when necessary, and remove them when they are no longer needed.
Mutations can affect information at different points in the pathway
A mutation is a change in a DNA sequence. Its consequences depend on where the change occurs and how it affects gene function.
In a protein-coding region, a single nucleotide substitution can produce a missense mutation, in which one amino acid is replaced by another. A substitution can also create a nonsense mutation, introducing a premature stop signal. Insertions or deletions that are not multiples of three can shift the reading frame, potentially altering every codon after the mutation.
But mutations do not have to change the amino acid sequence to matter. A DNA change can affect a promoter or another regulatory sequence and alter how much RNA is produced. Changes affecting RNA splicing can cause the wrong portions of an RNA molecule to be included or excluded.
The ultimate effect can range from essentially none to substantial disruption of cellular function. The effect depends on the gene, the particular change, the resulting molecular consequences, and the biological context.
Why this pathway matters
The DNA-to-RNA-to-protein pathway connects information with cellular activity.
DNA provides a stable store of instructions. Transcription selects and copies information into RNA. RNA processing prepares many transcripts for use. Ribosomes translate the information into amino acid sequences, and those sequences give rise to proteins with specific structures and functions.
The pathway also explains why a change in DNA can sometimes produce effects at the level of an entire organism. A genetic alteration can influence an RNA molecule, which can change a protein’s amount, sequence, structure, location, or activity. Because proteins participate in nearly every major cellular process, those molecular changes can ultimately affect how cells develop, communicate, metabolize nutrients, respond to their surroundings, or maintain themselves.
At the same time, the pathway is not a one-way story from a gene to a predetermined outcome. Cells constantly regulate which genetic information is used and how much of each product is made. The genome provides the information, but cellular machinery determines how that information is interpreted in a particular cell, at a particular time, under particular conditions.
That regulated movement of information—from DNA through RNA to functional molecules—is one of the central organizing principles of molecular biology.

