Scientists can reconstruct evolutionary history by comparing DNA sequences from different organisms. The basic idea is straightforward: species that inherited DNA from a more recent common ancestor tend to have more similar genetic sequences than species whose common ancestor lived farther in the past.
But evolutionary DNA analysis is more than simply counting matching letters. Scientists must decide which parts of genomes to compare, distinguish inherited similarities from similarities that arose independently, account for changes that have accumulated over time, and test whether different lines of evidence support the same evolutionary history.
The result is a phylogeny—a scientific hypothesis about the relationships among organisms.
DNA preserves evidence of common ancestry
DNA is made from four chemical bases, usually represented by the letters A, C, G, and T. The order of those bases encodes genetic information.
Mutations can alter DNA over generations. Most mutations have little or no effect on an organism’s survival and reproduction, although some are harmful and others can be beneficial. Once a mutation becomes established in a population, it can be inherited by later generations.
Suppose two species inherited a stretch of DNA from a common ancestor. Over time, mutations accumulated separately in the two descendant lineages. Their sequences would gradually become different. The more time that has passed, on average, the more differences may accumulate.
That pattern gives scientists a way to infer relationships. If two species have very similar versions of a particular DNA sequence, that similarity can be evidence that they share a relatively recent common ancestor. If their sequences are much more different, their shared ancestor may have lived farther back in evolutionary time.
The key phrase is common ancestry. DNA comparison does not require scientists to observe an evolutionary event directly. Instead, they use patterns of inherited similarities and differences to reconstruct the history that best explains the genetic evidence.
What scientists actually compare
Scientists rarely compare entire genomes indiscriminately. They often focus on corresponding regions of DNA, called homologous sequences, that can be traced to the same ancestral sequence.
For example, researchers might compare a particular gene across humans, chimpanzees, gorillas, and other primates. They align the sequences so that corresponding positions can be examined side by side.
A simplified alignment might look like this:
Species A ACGTACCTGAA
Species B ACGTACCTGGA
Species C ACGTTCCTGGA
Species D TCGTTCCTGGAThe differences are potentially informative. If a particular change is shared by two species but absent from others, it may indicate that those two species inherited the change from a common ancestor.
Insertions and deletions can also provide useful evidence. When DNA sequences contain extra or missing sections, computational methods can account for these gaps while determining which portions of the sequences correspond to one another.
The quality of this sequence alignment matters. If unrelated positions are mistakenly treated as equivalent, the resulting evolutionary analysis can be misleading.
Similar DNA does not automatically mean close relationship
One of the most important complications is that genetic similarity has several possible explanations.
Two organisms may have similar DNA because they inherited it from a common ancestor. This is the kind of similarity evolutionary biologists often want to detect.
But similar traits or sequences can also arise independently. Convergent evolution occurs when unrelated organisms independently evolve similar characteristics, often because they face similar environmental challenges. The underlying genetic similarities in such cases do not necessarily indicate a recent common ancestor.
Genes can also be copied within genomes. The resulting sequences, known as gene duplicates, can have different evolutionary histories. If scientists compare the wrong copies, they can reconstruct relationships incorrectly.
For this reason, researchers identify sequences carefully before using them to infer species relationships.
From DNA differences to an evolutionary tree
Once DNA sequences have been aligned, scientists use statistical and computational methods to infer a phylogenetic tree.
A phylogenetic tree is not necessarily a literal picture of how organisms look or how complex they are. It is a branching model of evolutionary relationships. Each branch represents a lineage, and points where branches meet represent inferred common ancestors.
Consider four species. If species A and B share a relatively recent common ancestor, while species C and D form another closely related pair, a simplified tree might represent the relationships as:
┌── A
┌───┤
│ └── B
─────────┤
│ ┌── C
└───┤
└── DThe important information is the branching pattern. A and B are more closely related to each other than either is to C or D in this particular hypothesis.
Scientists can construct trees in several ways. Some methods search for the tree that requires the fewest evolutionary changes under a particular model. Others use explicit probability models to estimate which evolutionary history is most consistent with the observed DNA sequences.
Maximum likelihood methods, for example, evaluate how probable the observed sequences would be under different possible trees and models of DNA change. Bayesian methods combine a model of sequence evolution with prior assumptions to estimate probabilities for possible trees.
Modern analyses commonly use computational programs because even a modest number of species and DNA positions can produce an enormous number of possible evolutionary trees.
Why scientists use models of DNA evolution
A simple count of differences can be useful, but it does not tell the whole story.
Imagine that an ancestral DNA position contained A. One descendant lineage changes A to G. Later, another mutation changes G back to A. If scientists examine only the present-day sequences, they may not realize that two mutations occurred.
Multiple substitutions can therefore accumulate at the same DNA position. Over long periods, the observed number of differences can underestimate the actual number of evolutionary changes.
Evolutionary models attempt to account for this. Different models make different assumptions about how frequently particular substitutions occur and whether some DNA positions change faster than others.
This matters especially when comparing organisms separated by long evolutionary periods. The farther apart two lineages are, the more opportunities there have been for mutations to overwrite earlier changes.
Not every part of DNA evolves at the same rate
Different regions of the genome can provide different kinds of evolutionary information.
Some DNA sequences change relatively quickly. These regions can be useful for distinguishing closely related populations or species because recent evolutionary differences are more likely to remain visible.
Other regions change slowly because their functions impose strong constraints. These conserved sequences can be useful for examining relationships among more distantly related organisms.
Genes themselves can contain regions under different evolutionary pressures. A mutation that changes an essential protein may be strongly selected against, while a mutation in a less constrained region may persist more readily.
Scientists therefore choose genetic regions appropriate to the evolutionary question they are asking rather than assuming that every DNA sequence carries the same historical signal.
Scientists compare many genes, not just one
A single gene can tell an incomplete story.
Genes have their own evolutionary histories, and those histories do not always perfectly match the history of the species carrying them. For example, ancestral populations can contain multiple genetic variants, and different descendant species may inherit different variants. This phenomenon, called incomplete lineage sorting, can produce gene trees that differ from the overall species tree.
Genes can also move between lineages in ways that complicate reconstruction. Horizontal gene transfer, particularly important in many microorganisms, allows genetic material to move between organisms rather than being passed only from parent to offspring.
Because of these complications, scientists often compare multiple genes or large portions of genomes. When independent genetic regions produce consistent patterns, confidence in the inferred evolutionary relationships generally increases.
Genome-scale data have made it possible to examine thousands or millions of DNA positions rather than relying on a small number of genes.
Mutations can help identify evolutionary branches
Shared genetic changes can be particularly informative.
Suppose a mutation appears in an ancestral population and is inherited by two descendant species. If closely related species outside that group do not possess the mutation, the shared change can provide evidence that the two species belong to the same evolutionary branch.
Scientists look for patterns like these across many positions in DNA. No single mutation necessarily settles a relationship. Instead, researchers evaluate the combined pattern of similarities and differences.
The same principle helps scientists distinguish ancestral characteristics from derived ones. An ancestral state was present earlier in the relevant lineage; a derived state arose later. Determining which is which often requires comparisons with additional species, called outgroups, that lie outside the group being studied.
DNA can be combined with fossils and anatomy
Genetic evidence is powerful, but scientists do not have to choose between DNA and other evidence.
Fossils provide physical evidence of organisms from different points in Earth’s history. Anatomy reveals patterns of structures that organisms share because of common ancestry. Embryological development, behavior, geography, and other evidence can also contribute to evolutionary reconstruction.
DNA is especially valuable because it contains information that can be compared quantitatively across living organisms and, in some cases, preserved biological material.
Fossils provide something DNA alone cannot: a direct record of organisms that lived at particular times in the past. Fossils can therefore help establish when evolutionary changes occurred and can constrain the timing of divergences inferred from genetic data.
When genetic, anatomical, and fossil evidence point toward compatible relationships, they provide mutually reinforcing lines of evidence.
Scientists can sometimes estimate when lineages diverged
DNA comparisons can also contribute to estimates of evolutionary timing.
The molecular clock is the general idea that genetic changes can provide a measure of elapsed evolutionary time. If researchers know, or can estimate, the rate at which particular genetic changes accumulate and have some independently dated evolutionary events to calibrate the analysis, they can estimate when lineages diverged.
The process is more complicated than assuming that mutations occur at one perfectly constant rate. Rates can differ among genes, lineages, and types of genetic change. Scientists therefore use models that allow for variation and use evidence such as well-dated fossils to calibrate molecular-clock analyses.
The resulting dates are estimates with uncertainty, not timestamps encoded precisely in DNA.
Ancient DNA adds another kind of evidence
DNA does not always have to come from living organisms.
Under suitable conditions, genetic material can survive in ancient remains. Ancient DNA has allowed researchers to compare extinct populations and species with their living relatives and, in some cases, to investigate relationships that would be difficult to resolve from anatomy alone.
Ancient DNA is particularly valuable because it can provide genetic information from an actual historical population rather than requiring scientists to infer everything from its modern descendants.
However, ancient DNA is fragile and can be contaminated by more recent biological material. Researchers must therefore use specialized laboratory procedures and computational methods to distinguish authentic ancient sequences from contamination and other sources of error.
How scientists know whether a tree is trustworthy
An evolutionary tree is a hypothesis supported by evidence, not an unquestionable photograph of the past.
Researchers test how strongly the data support different branches. One common approach is bootstrapping, in which the available sequence data are repeatedly resampled and trees are reconstructed from those resampled datasets. If a branch appears consistently, it receives stronger support than a branch that disappears easily when the data are changed.
Other statistical approaches provide probabilities or other measures of support.
Scientists also compare alternative models, examine different genes or genomic regions, and check whether conclusions depend heavily on particular assumptions. If different reasonable analyses repeatedly produce similar relationships, confidence in the result increases.
Disagreement among genes or datasets is not automatically a failure. It can reveal genuine biological processes, such as incomplete lineage sorting, gene duplication, or horizontal gene transfer. Understanding why datasets disagree can itself provide information about evolutionary history.
What DNA comparisons can—and cannot—tell us
DNA provides extraordinarily detailed evidence about biological relationships, but it does not record evolution as a simple chronological diary.
Scientists cannot generally look at two genomes and read off the exact sequence of every ancestral event. Some mutations leave no surviving trace, and different evolutionary histories can sometimes produce similar genetic patterns.
Instead, researchers infer the history that best explains the evidence while accounting for known biological processes and uncertainty.
That distinction is central to evolutionary genetics. The strength of DNA evidence comes not from a single dramatic genetic difference, but from the consistent patterns that emerge when many sequences, organisms, models, and independent lines of evidence are examined together.
By comparing inherited DNA changes across species, scientists can turn differences in modern genomes into evidence about the branching history of life—and continually refine that history as better genetic data and better methods become available.


