The Genetic Code: How Three Bases Specify One Amino Acid

The genetic code is the set of rules cells use to translate information in DNA and RNA into proteins. One of its most important features is that three RNA bases specify one amino acid. This three-base unit is called a codon.

Why three? Because RNA has four possible bases—adenine (A), uracil (U), cytosine (C), and guanine (G)—and cells need to specify 20 standard amino acids. A single base provides only four possible combinations, and two bases provide 16. Three bases provide 64 possible combinations, enough to encode all 20 amino acids, along with signals that tell the cell where to start and stop making a protein.

Understanding this simple numerical relationship helps explain both the elegance and the redundancy of the genetic code.

DNA stores information, but RNA carries the coding message

DNA contains the genetic instructions used by cells. For protein production, a stretch of DNA is first copied into messenger RNA (mRNA) through a process called transcription.

RNA uses four bases:

  • Adenine (A)
  • Uracil (U)
  • Cytosine (C)
  • Guanine (G)

The mRNA sequence is then read in groups of three bases. Each group is a codon, and each codon corresponds to a particular amino acid or to a signal involved in translation.

For example, an mRNA sequence might begin:

AUG GCU UUU GGA …

The ribosome reads these bases three at a time rather than one at a time. In this example, AUG corresponds to methionine, GCU to alanine, UUU to phenylalanine, and GGA to glycine.

The resulting amino acids are joined into a chain called a polypeptide. That chain folds into a functional protein.

Why does the code use three bases?

The answer comes from simple combinatorics.

There are four possible RNA bases at each position. If a codon contained only one base, there would be:

4 possible codons

With two bases, there would be:

4 × 4 = 16 possible codons

Neither number is enough to assign distinct signals to the 20 standard amino acids.

With three bases, however:

4 × 4 × 4 = 64 possible codons

That is more than enough. The genetic code therefore has 64 possible three-base combinations to represent 20 amino acids plus translation signals.

This does not mean every amino acid has only one codon. In fact, most have several. The 64 codons are distributed among the amino acids in a way that makes the code redundant, or degenerate: different codons can specify the same amino acid.

What exactly is a codon?

A codon is a sequence of three nucleotides in messenger RNA that specifies an amino acid or provides a signal for translation.

Nucleotides are the building blocks of RNA and DNA. Each RNA nucleotide contains one of the four bases A, U, C, or G. A codon is therefore simply three RNA bases arranged in a particular order.

For instance:

UUU is a codon for phenylalanine.

Changing the order changes the meaning. UUC also specifies phenylalanine, while other three-base combinations specify different amino acids or translation signals.

The ribosome moves along an mRNA molecule codon by codon. Transfer RNAs, or tRNAs, help match each codon with the appropriate amino acid. A tRNA carries an amino acid at one end and contains an anticodon at the other end that can pair with the corresponding mRNA codon.

This matching process converts the sequence of nucleotides into the sequence of amino acids that forms a protein.

The genetic code has 64 codons but only 20 standard amino acids

The apparent mismatch—64 codons versus 20 amino acids—is a defining feature of the genetic code.

Several codons can specify the same amino acid. For example, four different codons specify glycine:

GGU, GGC, GGA, and GGG

Likewise, leucine is specified by six codons, while some amino acids are specified by only two.

Three of the 64 codons do not specify amino acids. UAA, UAG, and UGA are stop codons. They signal that the ribosome should end translation.

AUG has a special role. It specifies methionine and commonly serves as the start codon that establishes where translation begins on an mRNA molecule.

The code is therefore not simply a dictionary in which every three-base sequence names a different amino acid. It is a system containing amino-acid assignments as well as instructions for beginning and ending protein synthesis.

Why redundancy matters

The redundancy of the genetic code means that a change in DNA does not always change the resulting protein.

Suppose a DNA mutation ultimately changes an mRNA codon from GAA to GAG. Both codons specify glutamic acid. The nucleotide sequence has changed, but the amino acid inserted into the protein remains the same.

Such a mutation is called a synonymous mutation because it does not alter the encoded amino acid.

Other mutations can change one codon into one that specifies a different amino acid. These are missense mutations. A mutation can also create a premature stop codon, producing a nonsense mutation.

Redundancy therefore provides some protection against certain sequence changes, although synonymous mutations are not necessarily biologically irrelevant. Changes in RNA processing, stability, translation efficiency, or other features can sometimes matter even when the amino acid sequence stays the same.

The order of the bases matters

The genetic code depends not only on which bases are present but also on their order.

Consider a sequence divided into codons:

AUG-CCU-GAA

If the sequence is read from the correct starting point, it represents one particular amino acid sequence. But inserting or deleting a single nucleotide can shift the grouping of all subsequent bases:

AUG-CCU-GAA
becomes
AUG-CUC-…

after an insertion or deletion, depending on the exact change.

Such an insertion or deletion can cause a frameshift mutation. Because codons are read in consecutive groups of three, shifting the reading frame changes nearly every codon that follows the mutation.

This is why the three-base structure is not merely a convenient way to label amino acids. The grouping itself is fundamental to how the information is interpreted.

How the genetic code connects DNA to protein

The information flow can be summarized as a sequence of related steps.

A gene in DNA contains a nucleotide sequence. During transcription, the relevant DNA sequence is copied into RNA. The resulting mRNA carries codons. During translation, a ribosome reads those codons and coordinates the addition of amino acids to a growing polypeptide chain.

The crucial transition is from nucleotide language to amino-acid language.

DNA and RNA use sequences built from four kinds of bases. Proteins use sequences built from amino acids. The genetic code provides the mapping between these two systems.

A three-base codon is the basic unit of that mapping because four possible bases, taken three at a time, create 64 possible combinations.

Why three bases, rather than two or four?

The choice of three bases is not arbitrary.

A two-base code could make only 16 combinations, which is fewer than the 20 standard amino acids. A three-base code makes 64 combinations, providing enough possibilities while keeping the basic unit relatively small.

A four-base code would make 256 possible combinations, far more than needed to represent the standard amino acids and translation signals.

Three therefore provides the minimum codon length that can accommodate the biological information required by the standard genetic code.

The resulting excess of codons is useful because it allows multiple codons to represent the same amino acid.

The genetic code is nearly universal

The basic codon assignments are shared across almost all forms of life, which is one reason the genetic code is described as nearly universal.

There are exceptions. Some organisms and cellular structures, particularly certain mitochondria and other specialized systems, use variant genetic codes in which particular codons have different meanings.

These variations do not change the central principle: biological information is read as nucleotide triplets, with each triplet assigned a specific meaning within the genetic code used by that system.

The key idea

Three bases specify one amino acid because a four-letter nucleotide alphabet requires at least three positions to produce enough combinations to encode the 20 standard amino acids. Four possible bases taken three at a time produce 64 codons. Those codons include assignments for all 20 amino acids, with many amino acids represented by multiple codons, as well as signals that regulate the beginning and end of translation.

That simple three-base structure is the foundation of how nucleotide sequences become protein sequences. It explains why the genetic code is compact enough to work efficiently, redundant enough to tolerate some changes, and precise enough to turn a sequence of bases into a specific sequence of amino acids.

Looking For Something Else?