Protein Folding: How a Chain of Amino Acids Becomes Functional

A protein begins as a linear chain of amino acids. Yet most proteins do not perform their jobs as extended strands. They fold into specific three-dimensional structures that allow them to bind molecules, catalyze chemical reactions, transmit signals, provide mechanical support, or carry out other cellular tasks.

Protein folding is the process by which a newly made polypeptide chain adopts its functional three-dimensional structure. It is not simply a matter of a chain being bent into shape. Folding involves many physical interactions, occurs through intermediate states in some proteins, and is influenced by the surrounding cellular environment. In some cases, proteins also require assistance from specialized molecules called chaperones.

Understanding folding helps explain a central principle of molecular biology: a protein’s amino acid sequence contains information that strongly influences its three-dimensional structure, and that structure in turn helps determine what the protein can do.

The sequence is the starting point

Proteins are built from 20 common amino acids. Cells link these amino acids together through peptide bonds, producing a polypeptide chain. The particular order of amino acids is determined by genetic information and gives each protein its primary structure.

The chain has a chemical backbone and side chains that differ from one amino acid to another. Some side chains are electrically charged, some are polar, some are nonpolar, and some contain chemically distinctive groups. These differences determine how parts of the chain interact with one another and with water.

The amino acid sequence therefore does more than specify the ingredients of a protein. It creates a pattern of chemical properties along the chain. As the polypeptide explores different shapes, those properties influence which conformations are energetically favorable.

But the sequence does not usually act like a rigid blueprint that dictates every movement. Folding is better understood as a physical process in which many possible conformations are explored, with interactions among the amino acids and the surrounding environment favoring some structures over others.

What actually drives a protein to fold?

A protein folds because certain molecular arrangements are more favorable than others under particular conditions. Several kinds of interactions contribute.

One of the most important effects involves hydrophobic amino acid side chains. These side chains tend to interact poorly with water. In an aqueous cellular environment, they often become buried inside a protein while more water-compatible groups remain exposed. This tendency, known as the hydrophobic effect, is a major contributor to protein folding.

Other interactions help stabilize the resulting structure. Oppositely charged side chains can attract one another. Polar groups can form hydrogen bonds. Some amino acids can participate in van der Waals interactions when atoms are packed closely together.

Cysteine introduces another important possibility. Two cysteine side chains can form a covalent disulfide bond under suitable conditions. These bonds can help stabilize certain proteins, particularly proteins that function in oxidizing environments such as the space outside many cells.

No single interaction normally explains a protein’s entire structure. Folding reflects the combined effects of many interactions, including interactions with water and, in living cells, with other molecules.

Folding creates structure at several levels

Protein structure is commonly described at four levels. They are not four completely separate stages; rather, they describe different aspects of the same molecular organization.

Primary structure

The primary structure is the amino acid sequence itself. It is the linear order of amino acids connected by peptide bonds.

A change in that sequence can sometimes have little effect, but in other cases a single amino acid substitution can alter folding, stability, or function. The effect depends on where the change occurs and what chemical properties are involved.

Secondary structure

Parts of the polypeptide backbone can form recurring local structures, especially alpha helices and beta sheets. These structures are stabilized largely by hydrogen bonds involving atoms of the protein backbone.

An alpha helix forms when a section of the chain coils into a regular spiral. A beta sheet forms when extended sections of a polypeptide align so that their backbones can hydrogen-bond with one another.

Secondary structure provides local organization, but it does not by itself determine the complete shape of a protein.

Tertiary structure

Tertiary structure is the overall three-dimensional arrangement of a single folded polypeptide chain.

At this level, distant portions of the amino acid sequence can come together. Hydrophobic residues may become buried, charged groups may form favorable interactions, and other side-chain contacts help pack the protein into a stable structure.

Many proteins have pockets, grooves, or surfaces created by this three-dimensional arrangement. These features can determine what molecules a protein can bind and what chemical reactions it can perform.

Quaternary structure

Some proteins consist of multiple polypeptide chains, or subunits. The way those subunits associate is called quaternary structure.

Hemoglobin, for example, is composed of multiple protein subunits that work together. Not every protein has quaternary structure; a protein made from a single polypeptide chain does not require it.

Folding is a search through many possible shapes

A polypeptide chain is flexible. Its bonds can rotate to some extent, allowing the chain to adopt a very large number of possible conformations.

If a protein had to examine every imaginable conformation one at a time before reaching its functional structure, folding would be impractically slow. Instead, folding generally proceeds through a biased physical process in which favorable interactions guide the chain toward particular regions of its possible conformational landscape.

This idea is often represented as an energy landscape. The unfolded chain occupies a high-energy and highly flexible collection of states. As it folds, interactions among its parts and with the environment tend to move it toward lower-energy states. The landscape can contain intermediate conformations and multiple pathways.

The lowest-energy structure is not necessarily the whole story. Proteins function in real cellular environments, where temperature, chemical conditions, binding partners, molecular crowding, and other factors can influence which conformations are populated. Some proteins also retain meaningful flexibility rather than existing as one perfectly rigid shape.

Proteins can fold while they are being made

Protein folding does not necessarily begin after the entire polypeptide has been synthesized.

Ribosomes build proteins by adding amino acids to the growing chain. As portions of the chain emerge from the ribosome, those regions may begin adopting local structures or interacting with other parts of the emerging protein.

This co-translational folding can affect the pathway the protein follows toward its final structure. The order and timing of a chain’s emergence can therefore matter, even though the completed amino acid sequence remains the fundamental source of structural information.

Once synthesis is complete, the protein may continue rearranging until it reaches its functional conformational ensemble.

Chaperones help proteins fold safely

Cells contain proteins called molecular chaperones that help other proteins reach or maintain appropriate conformations. Chaperones generally do not provide the final structural information themselves. Instead, they can influence the folding environment and reduce inappropriate interactions.

This matters because partially folded proteins can expose hydrophobic regions that normally become buried inside the finished structure. If exposed regions from different molecules interact, proteins can stick together and form aggregates.

Some chaperones temporarily bind exposed regions of proteins, giving them opportunities to fold without becoming trapped in unwanted interactions. Other chaperone systems provide specialized environments in which folding can occur.

Chaperones are therefore part of the cellular machinery that manages protein quality. They are especially important when cells experience conditions that increase protein misfolding, such as elevated temperature or other forms of stress.

Misfolding can change what a protein does

A protein does not have to be completely unfolded to be dysfunctional. A subtle structural change can alter an active site, disrupt a binding surface, or destabilize the protein enough that it is rapidly degraded.

Cells have quality-control systems that recognize and deal with many improperly folded proteins. Depending on the protein and the cellular circumstances, a misfolded molecule may be refolded, chemically modified, transported, or broken down.

A particularly difficult problem occurs when misfolded proteins aggregate. Aggregation can produce structures that are difficult for cells to remove and can interfere with normal cellular processes.

Some human diseases are associated with abnormal protein folding or aggregation. These include several neurodegenerative disorders, although the precise molecular mechanisms differ among diseases. Protein misfolding is therefore not merely a theoretical problem in biochemistry; maintaining protein structure is an important part of cellular health.

Folding and function are tightly connected

A protein’s function often depends on its three-dimensional structure.

An enzyme, for example, needs an appropriately shaped active site, where particular chemical groups are positioned so that a reaction can occur. A receptor needs a structure that allows it to recognize particular signaling molecules. Structural proteins need arrangements that provide appropriate strength, flexibility, or elasticity.

This is why changing an amino acid sequence can sometimes affect function even when the altered residue is not directly involved in the protein’s active site. The substitution may change the protein’s stability or subtly alter its overall shape.

The relationship also works in the other direction: binding another molecule can sometimes change a protein’s shape. Such conformational changes are common in biology. A protein is therefore not always a static object frozen into one structure. Its ability to shift between conformations can itself be essential to its function.

Not every protein follows the same folding path

There is no single universal folding mechanism.

Small proteins can sometimes fold rapidly and relatively simply, while larger proteins may contain multiple structural domains. A domain is a region of a protein that can often form a relatively independent structural and functional unit.

Some proteins fold efficiently on their own. Others depend heavily on chaperones or other cellular factors. Some proteins can adopt more than one biologically meaningful conformation, while certain proteins are intrinsically disordered in all or part of their normal functional state.

That last point is important: functional does not always mean rigidly folded. Some proteins contain flexible or disordered regions that become structured only when they interact with another molecule. This flexibility can be useful for signaling and regulation.

Why predicting protein structure is difficult

The amino acid sequence contains substantial information about a protein’s structure, but translating sequence into a reliable three-dimensional structure is a difficult computational problem.

The challenge comes partly from the enormous number of possible conformations and the complex interactions among residues. Even when the overall fold is known, predicting how a protein behaves dynamically, how it interacts with other molecules, or how mutations affect it can be considerably harder.

Modern computational methods have made major advances in predicting protein structures from amino acid sequences. These approaches have transformed structural biology, but a predicted structure is not identical to a complete understanding of a protein. Biological function can depend on molecular motion, chemical modifications, interactions with other proteins, membranes, nucleic acids, and the cellular environment.

From sequence to function

Protein folding can be viewed as a chain of connected ideas rather than a simple one-step transformation:

amino acid sequence → physical interactions → three-dimensional structure → molecular interactions → biological function

The sequence establishes the chemical possibilities. Physical forces and the cellular environment guide the chain toward particular conformations. The resulting structure creates surfaces, pockets, and moving parts that allow the protein to interact with other molecules and perform its role.

When folding goes wrong, those relationships can break down. When folding succeeds, a seemingly simple linear chain becomes a highly organized molecular machine—sometimes rigid in some regions, flexible in others, and capable of carrying out remarkably specific chemistry inside a living cell.

Looking For Something Else?