Coding vs. Noncoding DNA: What Do the Different Regions Do?

DNA is often described as the instruction manual for life, but that description can be misleading if it suggests that every part of DNA directly encodes a protein. In reality, only a fraction of the human genome consists of protein-coding sequences. Much of the rest is involved in regulating genes, producing functional RNA, organizing chromosomes, or performing other roles that scientists are still working to understand.

The distinction between coding DNA and noncoding DNA is therefore useful, but it is not the same as “important DNA” versus “useless DNA.” A region can be noncoding and still have an essential biological function.

Understanding the difference starts with how cells use DNA.

What coding DNA means

Coding DNA refers to DNA sequences whose information is used to specify the amino-acid sequence of a protein. Proteins carry out an enormous range of functions in cells, from catalyzing chemical reactions to forming structures, transporting molecules, and transmitting signals.

The protein-making process begins when a gene is transcribed: a cellular enzyme copies information from DNA into an RNA molecule. For a protein-coding gene, the resulting messenger RNA, or mRNA, is processed and ultimately used by ribosomes to assemble a protein.

The parts of a protein-coding gene that actually determine the protein’s amino-acid sequence are called coding sequences. In eukaryotic organisms such as humans, these sequences are generally found within exons, although not every portion of an exon necessarily codes for protein.

The genetic code connects three-nucleotide units called codons with amino acids. As a result, the sequence of coding DNA determines the order of amino acids in the resulting protein.

For example, changing a single DNA base can sometimes change a codon so that a different amino acid is inserted into the protein. Other changes can introduce a premature stop signal or alter how the RNA is processed. Such changes can affect the protein’s structure or function.

But protein-coding DNA is only one component of a gene.

Noncoding DNA does much more than “nothing”

Noncoding DNA is DNA that does not directly specify a protein’s amino-acid sequence. That broad category contains many very different kinds of sequences.

Some noncoding regions produce functional RNA molecules. Others regulate when, where, and how strongly genes are expressed. Some contribute to chromosome structure or help chromosomes behave properly during cell division. Still others have functions that remain incompletely understood.

The word noncoding therefore describes what a sequence does not do—directly encode a protein—rather than establishing that the sequence has no biological function.

This distinction matters because gene activity depends heavily on information outside protein-coding sequences.

Regulatory DNA controls when genes are used

A cell does not simply turn every gene on all the time. Different cell types activate different sets of genes, and the same cell can change gene activity in response to developmental signals or changes in its environment.

Much of this control comes from regulatory DNA.

Promoters help initiate transcription

A promoter is a DNA region associated with the beginning of a gene. Proteins involved in transcription recognize features of the promoter and help determine whether transcription begins.

Promoters are therefore part of the control system that allows a cell to use a gene.

Enhancers can increase gene activity

Enhancers are regulatory DNA sequences that can increase transcription of particular genes when the appropriate regulatory proteins bind to them.

Enhancers can sometimes be located far from the gene they regulate in the linear DNA sequence. DNA folding within the chromosome can bring the enhancer and gene into physical proximity, allowing regulatory proteins and the transcription machinery to interact.

Silencers can reduce gene activity

Some regulatory regions act to reduce or repress gene expression. These sequences, often called silencers, can recruit regulatory proteins that make transcription less likely or less active.

The result is a regulatory system in which genes are controlled rather than simply switched permanently on or off.

Introns are transcribed but usually do not encode protein

Many human genes contain both exons and introns.

When a protein-coding gene is transcribed, the initial RNA copy generally contains sequences corresponding to both. During RNA splicing, much of the intron sequence is removed, while exons are joined together to form the mature mRNA.

Introns are therefore part of genes and are transcribed, but they generally do not contribute directly to the amino-acid sequence of the protein.

This is one reason why “coding DNA” and “DNA inside a gene” are not interchangeable concepts. A gene can contain both protein-coding and noncoding portions.

Splicing can also be regulated in different ways. Through alternative splicing, cells can combine certain exons differently, allowing a single gene to contribute to multiple RNA or protein products.

Some noncoding DNA produces functional RNA

Not all useful genetic information ends up as a protein.

Some genes produce RNA molecules that perform functions directly. These are broadly called noncoding RNAs because the RNA itself is the functional product rather than an intermediate used to make a protein.

Examples include:

  • Ribosomal RNA (rRNA), which forms an essential part of ribosomes and participates in protein synthesis.
  • Transfer RNA (tRNA), which helps deliver amino acids during protein synthesis.
  • MicroRNAs (miRNAs), which can regulate gene expression by interacting with specific messenger RNAs.
  • Long noncoding RNAs (lncRNAs), a diverse group of longer RNA molecules with regulatory and other cellular functions.

These molecules demonstrate why RNA should not be thought of merely as a disposable copy of DNA instructions for proteins. In many cases, RNA is itself a functional molecule.

Some DNA helps organize chromosomes

DNA also has structural roles.

Telomeres are repetitive DNA sequences at the ends of chromosomes. They help protect chromosome ends from being mistaken for broken DNA and are associated with specialized proteins that maintain chromosome-end structure.

Centromeres are chromosome regions important for the accurate separation of chromosomes during cell division. They provide a platform for assembly of a protein structure that connects chromosomes to the machinery responsible for their movement.

Large portions of the genome are also packaged with proteins into chromatin. How DNA is packaged affects whether particular regions are accessible to the molecular machinery involved in gene regulation and transcription.

Thus, some genomic DNA contributes to the physical organization and behavior of chromosomes rather than encoding proteins.

Repetitive DNA makes up a large part of the genome

The human genome contains many sequences that occur repeatedly.

Some repetitive sequences are associated with recognizable biological functions, while others are derived from ancient mobile genetic elements and accumulated through evolutionary history. Transposable elements, for example, are DNA sequences capable of moving or copying themselves within genomes, although most such elements in the human genome are no longer capable of autonomous movement.

Repetitive DNA can contribute to chromosome structure, genome evolution, and regulation. At the same time, the biological significance of many individual repetitive sequences is not fully understood.

This is another reason it is inaccurate to divide the genome into two neat piles labeled “functional” and “nonfunctional.”

Coding and noncoding DNA work together

A protein-coding gene cannot be understood solely by looking at the sequence that determines its protein.

Its activity can depend on regulatory elements, chromatin structure, RNA processing, and signals from elsewhere in the cell. A mutation outside the protein-coding portion of a gene can therefore have a major effect if it disrupts an important regulatory sequence.

Consider two broad possibilities.

A mutation in a coding sequence might alter the protein itself—for example, by changing an amino acid or creating an abnormal stop signal.

A mutation in a regulatory region might leave the protein’s amino-acid sequence unchanged but cause the cell to make too much, too little, or none of that protein at the wrong time or in the wrong tissue.

Both kinds of changes can alter biological traits or contribute to disease.

“Junk DNA” is an outdated oversimplification

Scientists once had far less information about the functions of noncoding regions, leading to widespread use of the informal term “junk DNA.” The term can be misleading when applied broadly.

Some noncoding sequences clearly have important functions. Others may have effects that depend on cellular context. Still others may have little or no current function, or their function may simply not yet be known.

There is also an important distinction between biochemical activity and biological function. A DNA sequence may be copied into RNA or bind a protein without that activity necessarily being essential to the organism. Demonstrating that a genomic region has a reproducible, biologically meaningful function requires more than showing that something happens there.

For this reason, scientists generally describe specific genomic elements and their demonstrated roles rather than assuming that every noncoding sequence must serve an important purpose.

Coding DNA is not the same as “the important part”

The coding portion of a gene is crucial because it determines the sequence of its protein product. But the amount of protein-coding DNA in the human genome is relatively small compared with the genome as a whole.

The rest of the genome contains regulatory sequences, noncoding RNA genes, introns, repetitive elements, structural regions, and many sequences whose roles remain uncertain.

The more useful question is therefore not “Which DNA is useful: coding or noncoding?” It is “What role does this particular DNA sequence play?”

A protein-coding sequence can determine what a protein is. A promoter can influence whether the gene is transcribed. An enhancer can help determine where or when it is active. An intron can participate in the architecture and regulation of a gene and may contain regulatory information. A noncoding RNA gene can produce a molecule with its own function. A centromere can help a chromosome segregate properly.

Together, these different regions form a genome in which biological information is distributed across many kinds of DNA rather than concentrated exclusively in protein-coding sequences.

Looking For Something Else?