A decision tree is a machine learning model that makes predictions by following a sequence of questions about the data. Each question divides the data into smaller groups, and the model follows the branches that match the example it is evaluating until it reaches a final prediction.
This approach makes decision trees among the most intuitive machine learning methods. Rather than relying on a formula that is difficult to interpret, a tree represents its learned patterns as a series of decisions. It can predict whether an email is spam, estimate the price of a house, help assess the likelihood of equipment failure, or classify an image based on its measurable features.
Understanding how decision trees work requires looking at how they organize data, choose questions, learn from examples, and turn a path through the tree into a prediction. It also requires understanding their limitations, particularly their tendency to memorize training data when they grow too complex.
What a decision tree is and how it is organized
A decision tree is a model composed of connected points, called nodes, and the branches that link them. Each node represents a decision or a result, and each branch represents a possible outcome of a decision.
A typical decision tree has three main components. The root node is the starting point, where the model evaluates the first feature. An internal node contains a question or condition that divides the data. A leaf node, also called a terminal node, contains the final prediction.
Consider a model designed to predict whether a customer will renew a subscription. It might begin by asking whether the customer has used the service recently. If the answer is yes, the next question might concern the subscription plan. If the answer is no, the tree might examine how long the customer has been inactive. Each answer determines which branch the model follows.
The model continues through the tree until it reaches a leaf. That leaf provides the prediction, such as whether the customer is likely to renew or the estimated probability of renewal.
The features used in these questions are measurable properties of the examples being evaluated. They might include a customer’s recent activity, the length of a subscription, or the monthly price. In a medical dataset, features could include age, laboratory measurements, and reported symptoms. The appropriate features depend on the prediction task and the quality of the available data.
Although the tree’s structure resembles a human-made flowchart, a machine learning decision tree is usually constructed from examples rather than written as a fixed set of rules by a person. Its questions, thresholds, and branching structure are learned from data.
How a decision tree makes a prediction
Once a decision tree has been trained, making a prediction is generally straightforward. The model starts at the root node, evaluates the condition there, follows the corresponding branch, and repeats the process at each subsequent node.
Suppose a decision tree estimates whether a home is likely to sell for more than a chosen price. It might first ask whether the home’s size exceeds a particular threshold. Depending on the answer, it could examine the property’s age, neighborhood category, or another relevant feature. The model follows the resulting path until it reaches a leaf that assigns a prediction.
For a classification task, the leaf may specify a category, such as likely to sell above the target price or unlikely to do so. For a regression task, in which the goal is to predict a numerical value, the leaf may provide an estimated sale price.
A leaf can also represent a probability rather than only a category. For example, a classification tree might estimate a 70 percent probability of renewal for customers whose features lead to a particular leaf. That estimate is typically based on the outcomes of training examples that reached the leaf, possibly with adjustments or other estimation methods.
The distinction between training and prediction is important. During training, the algorithm searches for useful questions and builds the tree. During prediction, the trained structure is already fixed, so the model simply evaluates conditions along one path.
This separation makes prediction relatively efficient. A model does not need to compare a new example against every training example. It evaluates only the questions along the path from the root to the appropriate leaf.
How decision trees learn from data
Building a decision tree is a process of repeatedly dividing a training dataset into smaller groups. The central challenge is deciding which question should be asked at each node.
A training dataset contains examples with known outcomes. In a classification problem, those outcomes might be categories such as spam and not spam. In a regression problem, they might be numerical values such as house prices. The algorithm evaluates candidate splits to find questions that separate the examples in ways that help predict those outcomes.
For instance, an email dataset might include features such as the presence of certain words, the number of links, and information about the sender. The tree-building algorithm could test different conditions on these features and measure how well each condition separates spam from legitimate messages.
A split is useful when the resulting groups are more informative about the target outcome than the original group. If nearly all examples in one resulting group belong to the same category, that group is relatively easy to classify. If the groups contain a mixture of outcomes, the split may provide less useful information.
The algorithm selects a promising split, divides the data, and repeats the procedure within each resulting group. This process is called recursive partitioning. It continues until a stopping condition is reached, such as a limit on tree depth, a minimum number of examples in a node, or a point at which further splits are not considered worthwhile.
The resulting tree is not necessarily the best possible tree among all structures that could be built from the data. Finding a globally optimal decision tree can be computationally difficult, so common algorithms use greedy search: at each step, they choose the best split according to a local criterion. A locally good decision does not always produce the best overall tree, but this approach makes tree construction practical.
How a tree chooses its questions
Decision trees use mathematical measures to compare candidate splits. The specific measure depends on whether the model is performing classification or regression and on the algorithm being used.
For classification, two common measures are Gini impurity and entropy. Both describe how mixed the categories are within a group, although they calculate that mixture differently.
Gini impurity measures the likelihood that an example would be assigned an incorrect category if its label were selected at random according to the category proportions in the group. A group containing only one category has a Gini impurity of zero because there is no mixture of labels.
Entropy measures uncertainty in the distribution of categories. A group dominated by one category has lower entropy than a group in which the categories are evenly represented. Entropy is zero when every example belongs to the same category.
During training, a classification tree can compare the impurity of a parent node with the weighted impurity of its child nodes. The weights reflect how many training examples enter each child. A split that produces a substantial reduction in impurity is generally preferred over one that leaves the groups similarly mixed.
These measures are related but not identical, so they can sometimes favor different splits. Neither measure directly tells the model whether a feature is scientifically or causally important. They indicate how well a split separates the observed labels within the training data.
Regression trees use different criteria because their predictions are numerical rather than categorical. A common approach chooses splits that reduce the variation of target values within the resulting groups, often measured by the sum of squared errors or the corresponding mean squared error.
Imagine a dataset of homes with known sale prices. A split based on home size may separate relatively inexpensive homes from more expensive ones. If the prices within each resulting group are more tightly clustered, the split reduces prediction error on the training data and may be considered useful.
The exact rules vary among tree-building algorithms, but the underlying goal is consistent: find conditions that divide the available examples into groups whose outcomes are easier to predict.
How classification and regression trees differ
Decision trees can solve two broad kinds of prediction problems: classification and regression. The distinction lies primarily in what the model is asked to predict and how it evaluates potential splits.
A classification tree predicts a discrete category. Examples include identifying whether a transaction is potentially fraudulent, classifying an email as spam, or determining which species a biological sample most closely resembles based on its features.
The tree partitions the training examples according to their category labels. At a leaf, it may predict the most common category among the examples that reached that leaf. It can also provide class probabilities based on the distribution of labels there. These probabilities are estimates, not guarantees, and their reliability depends on factors such as sample size, model design, and how well the model generalizes.
A regression tree predicts a numerical quantity, such as a home’s sale price, the energy consumption of a building, or the expected duration of a process. Rather than separating examples by category, it groups examples whose numerical outcomes are relatively similar.
A regression leaf commonly predicts the mean of the target values among the training examples assigned to it. Some methods use a different representative value, such as the median, depending on the prediction objective and loss function.
This design has an important consequence: a standard regression tree typically produces a constant prediction within each leaf. If two homes follow the same path through the tree, they receive the same predicted price even if their actual prices differ. Predictions change when a feature crosses a split threshold and sends an example down a different path.
Both types of trees share the same general architecture and prediction process. Their main differences concern the target they predict, the criteria used to choose splits, and the values assigned to their leaves.
Why decision trees can be easy to interpret
One of the main attractions of decision trees is that their predictions can often be traced through a sequence of explicit conditions. A user can inspect the path taken by an example and identify which features influenced its result.
For a small tree, this can provide a clear explanation: the model predicted a particular category because the example satisfied a series of conditions. That transparency can help analysts identify unexpected rules, detect data problems, and communicate how a prediction was produced.
However, interpretability depends heavily on the tree’s size and structure. A tree with a few branches may be easy to understand, while one with hundreds or thousands of nodes can be difficult to follow. A technically explicit model is not automatically an understandable one.
A decision tree can also be misleading if its rules reflect quirks in the training data. A feature might appear influential because it happens to correlate with the target in the available sample, even when that relationship is unstable or has no causal significance.
For this reason, a tree’s explanation should be understood as a description of the model’s decision process, not necessarily an explanation of the real-world process that generated the data. If a tree uses income and location to predict a particular outcome, for example, its rules do not establish that either feature causes that outcome.
Interpretability is therefore a practical advantage, but it does not replace scientific scrutiny, domain knowledge, or evaluation on independent data.
Why decision trees can overfit
A decision tree that keeps splitting its training data can eventually create very small groups that closely match the examples it has already seen. This may improve performance on the training set while reducing performance on new data.
This problem is called overfitting. It occurs when a model learns not only the meaningful patterns in a dataset but also incidental details, random fluctuations, or noise that do not reliably recur.
Imagine a tree trained to predict equipment failure. A large tree might discover that a combination of unusual measurements, a particular maintenance date, and an uncommon operating condition corresponds to a handful of failures in the training data. Those conditions may have little predictive value in future cases. The model can fit the historical examples closely while making unreliable predictions on new equipment.
Several techniques help control this risk. One is to restrict the maximum depth of the tree, which limits how many successive decisions it can make. Another is to require a minimum number of training examples in a node before allowing further splits. A third is to limit the number of leaves or require a split to improve the model by a sufficient amount.
Pruning offers another approach. It removes branches that contribute too little to predictive performance. In cost-complexity pruning, for example, a training procedure balances the fit of the tree against a penalty for its complexity. The resulting model may sacrifice some accuracy on the training data in exchange for better generalization.
Choosing these controls requires evaluating the model on data that were not used to fit its rules. A validation set can help select the tree’s settings, while a separate test set can provide a more independent assessment of final performance. Cross-validation, which repeatedly trains and evaluates models on different partitions of the available data, is another common way to estimate how well a tree may generalize.
A smaller tree is not always better, and a larger tree is not always worse. The goal is to find a level of complexity that captures useful patterns without becoming overly dependent on the particular examples used for training.
Why small changes in data can change a tree
Decision trees have a notable weakness: their structure can be sensitive to changes in the training data. If two candidate splits perform almost equally well, a small change in the sample may cause the algorithm to choose a different one.
That first choice can influence every decision below it. Because the algorithm builds the tree recursively, a change near the root alters which examples reach each child node. Subsequent splits are then selected from different groups, potentially producing a substantially different tree.
This instability does not mean that every prediction from a tree is unreliable. Two trees with different structures may still make similar predictions on many examples. Nevertheless, structural differences can make an individual tree’s rules sensitive to sampling variation.
The issue is particularly important when training datasets are small, features are strongly correlated, or several candidate splits have similar predictive value. In such situations, the specific branch structure may reflect the accidents of the sample as much as stable relationships in the broader population.
One way to address this limitation is to combine many decision trees rather than rely on a single one. These methods are called ensemble methods because they aggregate predictions from multiple models.
In a random forest, many trees are trained using variations of the training data and feature selection. Their predictions are combined, often by voting for classification or averaging for regression. This generally reduces the influence of any one tree’s idiosyncrasies.
Gradient-boosted trees take a different approach. They build trees sequentially, with each new tree helping to correct errors or improve the objective left by the existing ensemble. Boosting can produce highly effective predictive models, although its behavior depends on the training procedure and the settings used.
Ensembles often improve predictive performance and stability, but they also tend to be harder to interpret than a single small tree. The choice between one tree and an ensemble therefore involves a trade-off between simplicity, interpretability, and predictive performance.
What decision trees do well and where they fall short
Decision trees can model nonlinear relationships and interactions between features without requiring the user to specify a particular mathematical formula in advance. A nonlinear relationship is one in which a change in a feature does not have a simple, constant effect on the outcome. An interaction occurs when the predictive relevance of one feature depends on the value of another.
For example, a tree can learn that a particular measurement matters only when another measurement exceeds a threshold. It can represent such conditional patterns through successive splits, making it useful for datasets in which the relationships among variables are not straightforward.
Trees can also handle mixtures of numerical and categorical features, provided the data are represented in a form supported by the algorithm. Their prediction procedure is relatively simple, and small trees can provide understandable decision rules.
Yet standard decision trees have limitations beyond overfitting and structural instability. Because each leaf often produces one fixed value or a category distribution, their predictions can change abruptly at split thresholds. A small difference between two otherwise similar examples may send them to different leaves and produce noticeably different predictions.
Trees can also struggle to represent smooth trends efficiently. A relationship that changes gradually may require many branches to approximate it, whereas another model may capture the same pattern with a more compact mathematical function. In regression, standard trees generally do not extrapolate trends naturally beyond the range of patterns represented in their leaves.
Performance also depends on data quality. Missing values, measurement errors, unrepresentative training examples, and biased labels can all affect the learned rules. The exact impact of missing data depends on the algorithm, since some tree implementations have explicit ways to handle missing values while others require preprocessing.
No model is universally best. A decision tree may be appropriate when transparent rules are valuable, the underlying relationships involve meaningful thresholds, or the task benefits from straightforward predictions. A random forest or boosted-tree model may be preferable when predictive accuracy is the primary goal and the added complexity is acceptable. Other model families may be better suited to particular kinds of data or scientific questions.
How to tell whether a decision tree is making reliable predictions
A decision tree’s ability to reproduce its training data is not sufficient evidence that it will work well in practice. The essential test is whether it predicts outcomes accurately for examples it did not use to learn its structure.
The appropriate evaluation measure depends on the task. Classification models may be assessed using accuracy, precision, recall, or measures that evaluate the balance between correctly identifying positive cases and avoiding false alarms. Regression models may be evaluated using mean absolute error, mean squared error, or root mean squared error, among other measures.
These metrics answer different questions. Accuracy measures the fraction of predictions that are correct, while precision considers how often positive predictions are correct and recall measures how many actual positive cases are identified. In a dataset where one category is much more common than another, accuracy alone can give an incomplete picture of performance.
For regression, mean absolute error measures the average magnitude of prediction errors, while squared-error measures penalize larger errors more heavily. The most appropriate metric depends on the costs of different mistakes and the purpose of the model.
Evaluation must also reflect the circumstances in which the model will be used. A tree trained on historical data from one population may perform differently on another population, particularly if the distribution of features or the relationship between those features and the target has changed. In time-dependent applications, testing on later observations can be more informative than randomly mixing earlier and later examples.
Reliable prediction also involves calibration when probabilities matter. A model is well calibrated when, among cases assigned a given probability, the observed outcome occurs at approximately that rate over an appropriate collection of cases. A tree that assigns a 70 percent probability to an outcome is not necessarily calibrated merely because 70 percent of the training examples in that leaf had the outcome.
Finally, good predictive performance does not establish causation. A decision tree can identify patterns that help forecast an outcome without determining why the outcome occurs. Establishing causal relationships requires additional assumptions, research designs, or evidence appropriate to the question.
Decision trees are useful because they turn patterns in data into a sequence of explicit decisions. Their predictions emerge from learned splits that progressively narrow the set of relevant examples, ending in a category, probability estimate, or numerical value. Understanding how those splits are chosen—and how complexity, data quality, and evaluation affect the result—makes it possible to use decision trees not just as convenient prediction tools, but as models whose strengths and limitations can be assessed critically.