The Central Dogma of Molecular Biology Explained

The central dogma of molecular biology describes the basic flow of genetic information in cells: DNA is used to make RNA, and RNA is used to make protein.

Written simply:

DNA → RNA → Protein

This idea is one of the foundations of modern biology. It explains how information stored in genes can ultimately produce the proteins that build cell structures, carry out chemical reactions, send signals, and perform countless other jobs.

The central dogma is useful because it separates three related but distinct molecules and processes. DNA serves as a long-term information store. RNA can carry or process genetic instructions. Proteins are major functional molecules that carry out much of the work inside cells.

The phrase “central dogma” does not mean that every piece of biological information must follow a single, unchanging route. Biology has important exceptions and additional pathways. Understanding what the central dogma actually says—and what it does not say—is therefore essential.

What the central dogma means

The central dogma concerns the transfer of sequence information between biological molecules.

DNA and RNA are nucleic acids made from chains of nucleotides. Proteins are chains of amino acids. The sequence of nucleotides in DNA can determine the sequence of nucleotides in an RNA molecule, and the sequence of an RNA molecule can determine the sequence of amino acids in a protein.

The three major information-transfer processes are:

  • Replication: DNA → DNA
  • Transcription: DNA → RNA
  • Translation: RNA → protein

Of these, transcription and translation form the familiar DNA → RNA → protein pathway.

Replication is different. It copies DNA so that genetic information can be preserved and passed to daughter cells. It does not convert DNA information into a protein.

The central dogma is therefore best understood as a framework for how sequence information can move between different types of biological molecules, rather than simply as a slogan about three molecules.

DNA: the long-term genetic information store

DNA, or deoxyribonucleic acid, contains the hereditary information of cells. In most organisms, DNA is organized into chromosomes, and specific stretches of DNA called genes contain instructions for producing functional biological products.

DNA has a double-stranded structure. The two strands are held together by predictable base pairing: adenine pairs with thymine, while cytosine pairs with guanine.

This complementarity is central to DNA’s role as an information store. Because the sequence on one strand determines the sequence on the other, a cell can use existing DNA as a template when making a new copy.

DNA replication

Before a cell divides, its DNA must be copied. During DNA replication, the two strands of the DNA double helix separate, and each strand serves as a template for building a complementary strand.

The result is two DNA molecules with essentially the same genetic information as the original, subject to copying errors and other sources of genetic variation.

Replication is not part of the DNA → RNA → protein conversion itself, but it explains how the genetic information underlying that pathway is maintained from one generation of cells to the next.

Transcription: DNA information becomes RNA

Transcription is the process of producing an RNA molecule using DNA as a template.

An enzyme called RNA polymerase moves along a region of DNA and builds an RNA strand whose nucleotide sequence is complementary to the DNA template strand. In RNA, the base uracil takes the place of thymine.

In a protein-coding gene, the resulting RNA can become messenger RNA (mRNA). The mRNA carries the information needed to specify a protein’s amino acid sequence from the relevant DNA region to the cellular machinery that performs translation.

In eukaryotic cells, which include human cells, an initial RNA transcript usually undergoes processing before becoming mature mRNA. This can include removal of noncoding segments called introns and joining of exons through a process called RNA splicing. The RNA may also receive other modifications that help it function properly.

Transcription is therefore more than simply “copying DNA.” It is a regulated process through which particular genetic information is selected and converted into an RNA molecule.

Translation: RNA information becomes protein

Translation is the process by which the nucleotide sequence of an mRNA is used to determine the amino acid sequence of a protein.

Translation takes place on ribosomes, molecular machines made from RNA and proteins. The ribosome reads the mRNA in groups of three nucleotides called codons.

Each codon generally specifies an amino acid or provides a signal involved in starting or stopping translation. Because there are four possible RNA nucleotides and codons contain three nucleotides, there are 64 possible codons. Several codons can specify the same amino acid, while some serve as stop signals.

Another type of RNA, transfer RNA (tRNA), helps connect codons with their corresponding amino acids. A tRNA carries a particular amino acid and contains an anticodon that can pair with a complementary codon in the mRNA.

As the ribosome moves along the mRNA, amino acids are linked into a growing chain. That chain is the protein’s polypeptide. After translation, the polypeptide folds and may undergo additional chemical modifications or processing to become a functional protein.

Why the sequence matters

The central dogma is fundamentally about information encoded in molecular sequences.

Consider a protein-coding DNA sequence. Its nucleotide sequence determines the sequence of the corresponding RNA. The RNA sequence, in turn, determines the order of amino acids incorporated into the protein.

A change in DNA can therefore sometimes change the resulting protein.

For example, a single nucleotide substitution in a protein-coding gene may change an mRNA codon. Depending on the specific change, the altered codon might still specify the same amino acid, specify a different amino acid, or become a signal that prematurely ends translation.

This illustrates why the central dogma is biologically important: changes in DNA sequence can propagate through molecular information-transfer steps and affect the structure or function of cellular products.

However, the relationship is not always straightforward. A gene can produce multiple RNA molecules through processes such as alternative splicing, and many RNAs do not encode proteins at all.

Not all RNA becomes protein

One of the most important qualifications to the simplified DNA → RNA → protein diagram is that RNA has many functions besides serving as a template for protein production.

Some RNA molecules are structural or functional components of cellular machinery. Ribosomal RNA is a major component of ribosomes, while transfer RNA helps deliver amino acids during translation.

Other RNAs participate in regulating gene expression, processing RNA, or carrying out specialized cellular functions. These molecules are examples of noncoding RNA: RNA that functions without being translated into a protein.

This does not contradict the central dogma. Rather, it shows why the simple three-word pathway should not be interpreted as “all RNA exists only to make proteins.”

The role of gene expression

The central dogma is closely connected to gene expression, the process by which information in a gene is used to produce a functional product.

A cell does not express every gene at the same time or to the same degree. Different cell types can contain essentially the same genome while producing different sets and amounts of RNAs and proteins.

Regulation can occur at many stages, including transcription, RNA processing, RNA stability, translation, and protein modification or degradation.

This regulation is a major reason cells with the same DNA can behave so differently. A muscle cell and a neuron, for example, use different subsets of their genetic information and therefore produce different patterns of RNAs and proteins.

The central dogma provides the information-flow framework, while gene regulation determines when, where, and how much of particular products are made.

Where reverse transcription fits

The central dogma is sometimes presented too rigidly, as though information can move only from DNA to RNA to protein. That interpretation is incorrect.

One important example is reverse transcription, in which RNA is used as a template to produce DNA:

RNA → DNA

An enzyme called reverse transcriptase carries out this process. Retroviruses use reverse transcription as part of their replication cycle, and reverse-transcription mechanisms also occur in other biological contexts.

Reverse transcription does not overturn the central dogma. The central dogma does not prohibit every possible transfer between DNA and RNA. Rather, it distinguishes information transfers involving nucleic acids from the transfer of sequence information from nucleic acid to protein.

Why protein does not normally transmit its sequence information back to nucleic acids

The central dogma makes a particularly important distinction involving proteins. Once nucleotide information has been used to specify an amino acid sequence, there is no ordinary cellular process that takes a protein’s amino acid sequence and directly converts it back into the corresponding DNA or RNA sequence.

This matters because the relationship between DNA and protein is not simply reversible.

A protein’s amino acid sequence can be influenced by a gene’s nucleotide sequence, but the amino acid sequence does not provide a unique way to reconstruct the original DNA sequence. The genetic code is degenerate, meaning that different codons can specify the same amino acid.

For example, knowing that a protein contains a particular amino acid does not necessarily reveal which codon encoded it.

The genetic code connects RNA to protein

The genetic code is the set of rules that relates mRNA codons to amino acids.

Because codons contain three nucleotides, an mRNA can be read as a series of three-letter units. Translation begins at an appropriate start signal and proceeds through the coding sequence until a stop codon is encountered.

The genetic code is described as degenerate because most amino acids are specified by more than one codon. It is also often described as nearly universal because the same basic code is used across a very broad range of organisms, although some organisms and cellular compartments have variations.

The code is what makes the final step of the central dogma possible: it provides the correspondence between a nucleic-acid sequence and a protein sequence.

The central dogma is not a complete description of molecular biology

The DNA → RNA → protein pathway captures a fundamental principle, but it leaves out much of what happens in real cells.

DNA is packaged and chemically modified. Genes are regulated. RNA is processed, transported, modified, degraded, and sometimes used directly as a functional molecule. Proteins fold into three-dimensional structures, interact with other molecules, undergo chemical modifications, and are eventually broken down.

There are also biological information flows and molecular processes that do not fit neatly into the simplified diagram. RNA can be copied from RNA in certain biological systems, and RNA can be converted into DNA through reverse transcription.

These processes do not make the central dogma useless. They show why its precise meaning is more informative than the oversimplified version often taught first.

A useful way to remember the pathway

For a protein-coding gene, the basic sequence is:

DNA → transcription → RNA → translation → protein

DNA provides the original nucleotide sequence. Transcription produces an RNA copy or transcript of the relevant information. Translation reads the appropriate RNA sequence and uses the genetic code to assemble a chain of amino acids.

Replication operates alongside this pathway:

DNA → replication → DNA

Its purpose is to preserve genetic information rather than produce a protein.

The central dogma therefore gives molecular biology a simple organizing principle: genetic sequence information can be stored in DNA, transferred to RNA, and used to specify proteins, while RNA itself can also serve important functions independent of protein production.

Looking For Something Else?