Machine learning models fail when they do not capture the patterns needed to make reliable predictions on new data. Two fundamental reasons are overfitting and underfitting. An overfitted model learns the training data too closely, including its noise and accidental quirks. An underfitted model fails to learn the underlying relationships in the first place.
These problems reflect a central challenge in machine learning: finding a model that is complex enough to recognize meaningful patterns but not so complex that it mistakes random variation for useful information.
Understanding overfitting and underfitting helps explain why a model can perform impressively during development yet make poor predictions in practice, why more training does not always improve results, and how researchers and engineers can build systems that generalize beyond the examples they have already seen.
What overfitting and underfitting mean
Machine learning is a process in which a computer system learns patterns from data rather than relying entirely on rules written by a programmer. During training, an algorithm adjusts a model’s internal parameters to reduce errors on a collection of examples.
For instance, a model designed to predict home prices might learn relationships between sale prices, floor area, location, property age, and other characteristics. Once trained, it uses those relationships to estimate the price of a home it has not encountered before.
The important question is not simply whether the model can reproduce the examples used for training. It is whether the model can make accurate predictions on new examples drawn from the situations in which it will actually be used. This ability is called generalization.
Overfitting and underfitting represent two different failures of generalization. An underfitted model is too limited to capture important relationships in its data. An overfitted model captures not only genuine relationships but also incidental details that do not reliably recur. Both can produce inaccurate predictions, but they do so for different reasons.
Consider a model that estimates a person’s commute time from distance traveled. If it assumes that every additional mile adds exactly the same amount of travel time, it may overlook important factors such as traffic, road type, and departure time. Its assumptions may be too simple for the task, producing underfitting.
At the other extreme, a model might memorize the commute times associated with individual routes and particular days. It could perform exceptionally well on previously recorded journeys but poorly when predicting a new trip or an unfamiliar traffic situation. This is overfitting.
The goal is not to make a model as simple or as complicated as possible. It is to find a level of complexity that captures the relationships that matter without becoming dependent on details that will not reliably predict future outcomes.
How underfitting occurs
Underfitting happens when a model cannot represent the important patterns in the data, even after it has been trained adequately. Its predictions remain systematically inaccurate because its assumptions, structure, or learned parameters are insufficient for the problem.
One cause is excessive simplicity. A linear model, for example, represents relationships using a straight line or a linear combination of input variables. This can work well when the underlying relationship is approximately linear. But if the relationship changes sharply across different conditions, a purely linear model may not capture it adequately.
Imagine predicting electricity consumption from outdoor temperature. Heating demand may increase when temperatures fall, while air-conditioning demand may increase when temperatures rise. A model that assumes consumption changes in only one direction may fail to describe this relationship. A more suitable model could account for the different patterns at high and low temperatures.
Underfitting can also arise from poor input data. If a model is intended to predict traffic congestion but receives only the day of the week, it may lack essential information about accidents, weather, road closures, or current traffic conditions. Even a sophisticated algorithm cannot reliably recover information that its inputs do not contain.
Insufficient training can contribute as well. Some models need many examples or enough optimization steps to learn their parameters effectively. If training stops too early, the model may not have captured even the relationships it is capable of representing.
These causes require different remedies. A model with an inadequate structure may need a more flexible design. A model missing important information may need better input variables. A model that has not learned sufficiently may benefit from additional training. Simply increasing model complexity will not solve every form of underfitting.
An underfitted model typically performs poorly on both its training data and unseen data. Its training error remains relatively high because it has failed to explain the examples it was given. Its test error is also high because the underlying patterns remain poorly represented.
How overfitting occurs
Overfitting occurs when a model learns the training examples too specifically. Instead of identifying relationships that are likely to persist, it becomes sensitive to details that are peculiar to the particular data used during training.
Real-world data contain both meaningful signals and noise. A signal is a pattern that carries useful information about the outcome being predicted. Noise includes random fluctuations, measurement errors, inconsistencies, and other details that may not reflect a stable relationship.
A sufficiently flexible model can sometimes fit both. This may lower its training error, but learning noise does not necessarily improve its ability to predict new cases.
Consider a model trained to distinguish photographs of cats from photographs of dogs. Suppose most cat photographs in the training set have light-colored backgrounds, while most dog photographs have dark backgrounds. A model might learn to associate background color with the animal category instead of focusing on the animals’ features.
The model could achieve strong results on the training photographs because background color happens to correlate with the labels. Yet it might misclassify a cat photographed against a dark background or a dog photographed against a light one. The relationship it learned was real in the training data but unreliable as a general rule.
Overfitting does not require a model to memorize every training example exactly. It can also occur when a model learns complicated, unstable relationships that fit the available data unusually well. The essential problem is that its learned patterns do not transfer reliably to new examples.
Several conditions make overfitting more likely. A highly flexible model can represent a wide range of relationships, including accidental ones. A small or unrepresentative training set gives the model fewer examples from which to distinguish stable patterns from chance associations. Noisy labels, irrelevant input variables, and repeated experimentation against the same evaluation data can also encourage overfitting.
Training duration matters in some settings, too. As optimization proceeds, a model may first learn broad, useful relationships and later become increasingly adapted to peculiarities of the training examples. This pattern is common in certain learning problems, although it is not a universal rule. Some models continue to improve on unseen data for a long time, while others overfit quickly.
The characteristic sign of overfitting is a widening gap between performance on the training data and performance on genuinely unseen data. Training error becomes low, but test error remains higher than expected or begins to increase as training continues.
Why training accuracy can be misleading
A model’s performance on its training data reveals how well it fits the examples it has already seen. It does not, by itself, establish whether the model will work in the real world.
This distinction explains why machine learning systems can appear successful during development and disappoint after deployment. A model can learn details specific to its training set that have little value outside that set. Conversely, a model that has not learned enough may perform poorly even on familiar examples.
To assess generalization, practitioners typically divide available data into separate subsets. The training set is used to learn the model’s parameters. A validation set helps compare model designs, tune settings, and decide when to stop training. A test set is reserved for a final, relatively unbiased assessment of the chosen approach.
The distinction matters because repeated decisions based on an evaluation set can gradually adapt the development process to that set. If an engineer tests many model configurations and repeatedly chooses whichever performs best on the same validation data, the selected model may begin to benefit from chance characteristics of those examples. The validation results can then become more optimistic than performance on truly new data.
A separate test set helps reduce this problem, provided it remains independent of the model-selection process. It should not be repeatedly used to guide changes to the model. Once test results influence development decisions, the test set is no longer a fully independent final check.
Data splitting must also reflect how the model will be used. Randomly splitting records may be inappropriate when examples from the same person, household, hospital, or physical object appear multiple times. Related records can share information that makes evaluation artificially easy.
For example, a medical model evaluated on records from patients represented in both the training and test sets may perform differently from one evaluated on entirely new patients. Similarly, a model that predicts future demand should generally be evaluated using a time-based split, with earlier observations used for training and later observations reserved for evaluation. Training on information that would not have been available at prediction time can produce misleadingly strong results.
A reliable evaluation therefore requires more than separating rows in a dataset. It requires asking whether the test examples genuinely represent the uncertainty the model will face in practice.
The relationship between model complexity and generalization
Model complexity describes, broadly, how many different patterns a model can represent. It is not determined solely by the number of parameters. Architecture, regularization, training procedures, input features, and the nature of the data all influence the patterns a model can learn.
A simple model has limited flexibility. This can protect it from fitting random fluctuations, but it may be unable to capture important relationships. A highly flexible model can represent more complicated patterns, but it may also fit accidental features of the training data.
This creates a trade-off between two sources of predictive error. Bias refers to systematic error associated with a model’s assumptions or limitations. A model with high bias may consistently miss important patterns. Variance refers to how sensitive a model’s learned predictions are to changes in the training data. A model with high variance may learn substantially different patterns when trained on different samples from the same underlying population.
Underfitting is commonly associated with high bias, while overfitting is commonly associated with high variance. These are useful connections, not exact definitions: real models can exhibit several kinds of error at once, and their behavior depends on the task and data.
The classical bias–variance framework helps explain why adding flexibility can initially improve a model. A more expressive model may capture relationships that a simpler model misses, reducing systematic error. Beyond some point, however, additional flexibility may cause the model to respond too strongly to sampling noise, increasing variance.
In many traditional settings, this trade-off produces a pattern in which test error falls as complexity increases, reaches a favorable region, and then rises. The best model lies near the point where it captures meaningful structure without becoming excessively sensitive to the training sample.
However, this pattern is not universal. Modern machine learning, especially deep learning, can display more complicated relationships between model size and test performance. Under some conditions, increasing the number of parameters beyond a conventional optimum can improve generalization again, a phenomenon often called double descent. Such behavior does not eliminate overfitting; it shows that simple descriptions of the relationship between complexity and error do not capture every learning regime.
The practical lesson remains the same: model complexity should be judged by evidence from appropriate evaluation data, not by assumptions that simpler models always generalize better or that larger models are inherently more capable.
How to recognize overfitting and underfitting
The most useful diagnostic evidence comes from comparing training performance with validation performance over time and across candidate models.
An underfitted model generally has high training error and high validation error. It fails to explain the training examples adequately, and its performance on unseen data is correspondingly weak. If training and validation errors are both high, the model may be too limited, inadequately trained, or supplied with insufficient information.
An overfitted model often has low training error but substantially higher validation error. It has learned the training examples well without achieving comparable performance on new ones. If validation error improves initially and then worsens while training error continues to fall, overfitting is a plausible explanation.
A well-generalizing model achieves reasonably low error on both sets, with a gap that is acceptable for the task. The exact size of that gap depends on the problem, the amount of data, the evaluation metric, and the uncertainty inherent in the outcomes.
These patterns are diagnostic clues rather than definitive proof. High validation error can also result from a distribution mismatch, meaning the validation data differ systematically from the training data. A poor choice of metric, incorrect labels, data leakage, or a genuinely unpredictable outcome can also complicate interpretation.
The metric itself matters. Classification accuracy, for example, can conceal serious weaknesses when one category is much more common than another. A model that predicts the most frequent category every time may appear accurate while failing to identify the less common category that matters most. Depending on the application, precision, recall, calibration, or other measures may provide a more informative assessment.
Calibration describes whether predicted probabilities correspond to observed frequencies. A model that assigns a 70 percent probability to many comparable events is well calibrated when approximately 70 percent of those events occur. A model can rank outcomes effectively yet produce poorly calibrated probabilities, so predictive quality is not always captured by a single score.
The appropriate evaluation method should reflect the consequences of mistakes. In a medical screening system, missing a serious condition may be more costly than producing an unnecessary follow-up. In a recommendation system, other trade-offs may matter more. A model is not meaningfully successful simply because it performs well according to a metric that ignores its intended use.
How to reduce underfitting
Reducing underfitting begins with identifying what prevents the model from learning the necessary relationships.
One approach is to use a model capable of representing the patterns in the data. A linear model may be insufficient for a strongly nonlinear problem, while a more flexible method may capture the relationship more effectively. Additional input features can also help when they supply relevant information about the outcome. Transformations, interactions between variables, or representations learned from raw data may expose patterns that the original inputs conceal.
The quality of the training process matters as well. An optimization algorithm adjusts a model’s parameters to reduce a chosen objective, or mathematical measure of error. Poor settings, an unsuitable learning rate, or premature stopping can prevent effective learning. Improving the training procedure may allow a model to reach a better solution without changing its overall design.
More training data can help, but it is not always the answer to underfitting. If the model is too restricted to express the relevant relationship, additional examples will not remove that limitation. Likewise, adding more variables will not necessarily help if they are irrelevant or unreliable.
Underfitting can also reflect a mismatch between the objective and the real goal. A model trained to minimize average error may perform poorly on rare but important cases. Adjusting the objective or evaluation criteria may be necessary, provided the changes are justified by the application’s needs and do not merely optimize a convenient score.
The key is to determine whether the problem comes from limited model capacity, missing information, ineffective training, or a poorly chosen objective. Each calls for a different intervention.
How to reduce overfitting
Reducing overfitting means encouraging a model to learn relationships that are likely to remain useful beyond its training examples.
One common method is regularization, which discourages certain forms of model complexity. In many algorithms, regularization adds a penalty to the training objective when the model’s parameters become excessively large or complicated according to a specified measure. This encourages solutions that fit the data without relying as heavily on extreme parameter values. The exact effect depends on the method and model; regularization does not simply remove every complicated pattern.
Another approach is to limit the model’s effective flexibility. A decision tree, for example, can be restricted in depth or in the minimum number of examples required to form a leaf. These limits can prevent the tree from building branches around small, unrepresentative groups of observations. For other model families, different constraints or architectures may be appropriate.
Early stopping can help when a model begins to overfit as training continues. Performance on a validation set is monitored, and training is stopped when additional updates no longer improve the chosen measure. This method is useful when validation performance provides a meaningful indication of generalization, but it requires care because repeatedly tuning decisions against the same validation set can itself lead to overfitting.
Data augmentation can be valuable when the task allows realistic variations of existing examples. In image recognition, for instance, modest changes in orientation, cropping, or lighting may help a model learn features that are not tied to a particular presentation of an image. The transformations must preserve the relevant meaning. An alteration that changes a label or creates unrealistic examples can make learning worse rather than better.
Collecting more representative training data is another powerful strategy. Additional examples can expose the model to more of the natural variation it will encounter in practice, making it harder to rely on accidental patterns. Yet quantity alone is insufficient. A million nearly identical examples may provide less useful diversity than a smaller, carefully assembled dataset that better represents the intended population.
Feature selection and data cleaning can also reduce overfitting. Removing irrelevant or unreliable inputs may make it harder for a model to exploit spurious associations. Correcting labeling errors and addressing inconsistent measurement can improve the signal available during training. Still, automatically removing unusual observations is dangerous: rare cases may be legitimate and essential to the task.
No single technique guarantees success. Regularization that is too strong can push a model toward underfitting. Excessively aggressive data augmentation can distort the problem. Stopping training too early can prevent useful learning. Effective control of overfitting requires evaluating interventions on data that remain independent of the training process.
Why data quality and representativeness matter
Overfitting and underfitting are often discussed as problems of algorithms, but the data can be equally important. A model learns from the examples it receives, and those examples determine which relationships it can discover and which limitations it inherits.
A training set may be too small to represent the range of conditions that matter. It may omit certain populations, environments, outcomes, or unusual events. It may contain measurement errors or reflect historical decisions that should not be treated as objective truth. A flexible model can amplify these weaknesses, while a restricted model may fail to capture the differences that the data do contain.
Representativeness is especially important because the data encountered after deployment may differ from the training data. This is known as distribution shift. Changes in consumer behavior, medical practice, economic conditions, equipment, or data collection can alter the relationships a model relies on.
For example, a model trained to forecast product demand during a period of stable supply may become less reliable when supply constraints or customer preferences change. Even a model that was not overfitted during development can perform poorly under such conditions. Its failure may reflect a change in the environment rather than excessive learning of training noise.
This distinction has practical consequences. Regularization cannot reliably correct every form of distribution shift. A model may need updated data, revised input variables, a different target, or a redesigned evaluation process. Monitoring performance after deployment can reveal deterioration that a static test set could not anticipate.
Data leakage presents another danger. Leakage occurs when information unavailable at the time of a real prediction enters training or evaluation in a way that makes performance appear better than it should. For example, using a variable recorded only after a medical diagnosis to predict that diagnosis would create an unrealistic task. The model might perform well in evaluation but fail when asked to make predictions before the diagnosis occurs.
Preventing leakage requires understanding how data are generated, what information is available at prediction time, and how records are related. Good model design therefore depends on knowledge of the real-world process, not just proficiency with an algorithm.
Why the best model is not necessarily the most accurate on paper
A model’s apparent quality depends on how it is selected, tested, and used. The highest score in a development experiment may not correspond to the most dependable system in practice.
When many models are compared, some will perform well on a validation set partly by chance. The more alternatives researchers test against the same evaluation data, the greater the opportunity to select one that benefits from accidental features of that data. This is a form of selection bias: the chosen result may look unusually strong because it was selected from many attempts.
Cross-validation can help estimate generalization more robustly when data are limited. In this procedure, the dataset is divided into several parts, and models are trained and evaluated across different combinations of those parts. The resulting performance estimates provide a broader view than a single arbitrary split. However, cross-validation does not automatically solve leakage, distribution shift, or the problem of repeatedly selecting models against the same data. The splits must be appropriate to the task.
A final evaluation should be as independent as practical and should reflect the intended conditions of use. For a time-sensitive forecasting model, that may mean testing on later periods. For a model expected to serve new users, it may mean separating users rather than individual records. For a system deployed across different locations, it may mean checking whether performance holds across those locations.
The cost of errors should also influence model selection. A modest improvement in average accuracy may not justify a model that fails more often on a vulnerable group or becomes unreliable under important operating conditions. Overall performance can hide substantial differences across subgroups, so evaluation should examine the aspects of reliability that matter for the application.
In consequential settings, predictive performance is only one part of responsible deployment. Models may also require checks for fairness, robustness, privacy, interpretability, and operational safety. A model can generalize statistically and still be inappropriate for a particular use if its errors cause unacceptable harm.
Building models that learn the right lessons
Overfitting and underfitting reveal a fundamental limitation of learning from finite data: a model cannot directly observe every situation it will eventually encounter. It must infer broader relationships from a limited collection of examples, and those inferences are inevitably shaped by its assumptions, design, and training process.
An underfitted model does not learn enough of the relevant structure. An overfitted model learns details that are too specific to the examples it has seen. A model that generalizes well finds a useful balance, capturing patterns that persist while remaining appropriately insensitive to incidental variation.
Achieving that balance requires more than selecting an algorithm. It involves defining the prediction task clearly, collecting suitable data, choosing an appropriate model, evaluating it without leakage, monitoring performance on unseen examples, and revisiting assumptions when conditions change.
The most reliable machine learning system is therefore not the one that best explains its training data. It is the one whose learned relationships continue to support useful, well-calibrated predictions in the situations for which it was designed.