Artificial intelligence systems do more than recognize patterns or generate predictions. Many must also decide what to do next, determine which steps will achieve a goal, and adjust their behavior when circumstances change. This process, known as AI planning and decision-making, allows autonomous systems to move beyond responding to individual inputs and toward carrying out sequences of actions.
An autonomous vehicle, for example, must decide when to slow down, whether to change lanes, and how to respond when another driver behaves unexpectedly. A warehouse robot must select a route, avoid obstacles, and determine how to move objects efficiently. A software agent may need to gather information, use digital tools, and complete a multistep task without receiving instructions for every action.
In each case, the system faces a common challenge: it must choose an action based on what it knows, what it wants to accomplish, and what it expects to happen next.
AI planning and decision-making provide methods for solving this challenge. They combine representations of goals and actions with reasoning about possible outcomes, uncertainty, constraints, and the costs of different choices. The specific methods vary, but the underlying objective is consistent: select actions that help achieve a goal while accounting for the consequences.
What AI planning and decision-making mean
Planning and decision-making are closely related, but they address different parts of the problem.
AI planning focuses on identifying a sequence of actions that can transform an initial situation into a desired one. A planner might determine how a robot can move from one room to another, collect an object, and deliver it to a destination. It considers which actions are available, what conditions those actions require, and how they change the environment.
AI decision-making focuses on choosing among available actions. It considers which choice is preferable given the system’s objectives, constraints, information, and expectations about the future. A decision-making system might choose between a short route with unpredictable traffic and a longer route that is more reliable.
Planning typically emphasizes how to reach a goal. Decision-making emphasizes which action or strategy to choose. In practical systems, the two processes overlap: a planner generates possible courses of action, and a decision-making mechanism evaluates or selects among them.
Consider a home robot instructed to clean a room. It needs to identify which areas require cleaning, determine where its cleaning equipment is located, choose a route, and decide how to respond if furniture blocks the planned path. Planning organizes the task into steps, while decision-making helps select those steps and adapt them when conditions change.
Neither process necessarily requires human-like understanding. An AI system can select effective actions by using a mathematical model, a set of rules, a learned policy, or a combination of these methods. Its ability to act successfully depends on how well its internal representations and decision procedures match the task and environment.
How an AI system turns a goal into actions
Although different AI architectures work in different ways, many autonomous systems follow a recurring cycle: represent the current situation, identify the goal, generate possible actions, evaluate their consequences, select an action, and observe what happens.
The process begins with a state, a description of the situation relevant to the decision. For a mobile robot, the state might include its location, direction, battery level, nearby obstacles, and the position of its target. For a digital assistant, it might include the user’s request, the information already gathered, the available tools, and the task’s completion status.
A state representation does not have to include every detail of the world. It needs to contain enough information to support the decisions the system must make. Omitting an important variable, such as battery charge in a long-distance robot mission, can lead to a plan that appears workable but fails in practice.
Next, the system defines a goal or objective. A goal might be to reach a destination, complete a sequence of software operations, or deliver an object without damaging it. Some goals are expressed as conditions that must be satisfied. Others are represented by an objective function, a mathematical expression that assigns a value to possible outcomes.
The system then considers the actions available from its current state. Each action has requirements and consequences. A robot can move forward only if its path is sufficiently clear; a digital agent can retrieve a file only if it has access to the relevant system. The planner must account for these restrictions rather than simply select actions that sound useful.
Possible actions can be evaluated by how much progress they make toward the goal, how much time or energy they require, what risks they introduce, and whether they preserve future options. Once an action is selected, the system executes it and receives new information about the result.
This final step is essential. Real environments rarely behave exactly as expected. A robot may encounter an obstacle that was not detected earlier, or a software operation may fail because a resource is unavailable. The system must compare the observed result with its expectations and decide whether to continue, revise its plan, or pursue a different strategy.
This repeated process is often called a perception-action loop. Perception supplies information about the environment, decision-making selects an action, and the resulting action changes the situation the system must perceive next.
How AI planners evaluate possible futures
Planning becomes difficult when actions have consequences that extend beyond the immediate moment. A choice that looks beneficial now may create problems later, while a small initial expense may enable a much better outcome.
A planner must therefore reason about future states, not just immediate actions.
Suppose a warehouse robot needs to deliver a package. The shortest route may pass through a crowded aisle, while a longer route avoids congestion. A planner that considers only distance might choose the shorter path. One that accounts for expected delays, collision risk, and energy use may prefer the alternative.
To make such comparisons, planners use models of how actions change states. In a simplified environment, the model may be a set of explicit rules. In a more complex environment, it may be a mathematical description of movement, a learned prediction model, or a combination of both.
The planner uses this model to estimate the consequences of different action sequences. It may search through possible sequences until it finds one that satisfies the goal, or it may compare candidate plans according to a cost function.
A cost function assigns a numerical cost to a state, action, or sequence of actions. Depending on the task, it might represent travel time, fuel consumption, monetary expense, exposure to hazards, or a weighted combination of several factors. A planner can then seek a low-cost solution that satisfies the required constraints.
The objective is not always to minimize cost. Some systems maximize a reward, which assigns higher values to more desirable outcomes. Others seek to satisfy a goal while minimizing resource use, or they must meet several competing requirements.
These distinctions matter because the best action depends on how success is defined. A delivery system optimized only for speed may take unacceptable risks. A robot optimized only for conserving energy may fail to complete its mission promptly. An objective function makes priorities explicit, but it does not automatically guarantee that those priorities are appropriate.
Search and the problem of too many possibilities
Even a simple task can produce an enormous number of possible action sequences. If a robot has several movement choices at each step, the number of possible paths grows rapidly as the number of steps increases. More complicated tasks add choices about timing, resources, dependencies, and other agents.
One way to solve this problem is search. A search algorithm explores possible states or action sequences, using rules to determine which possibilities to examine and when to stop.
Breadth-first search explores states in order of the number of actions needed to reach them. Under appropriate conditions, it can find a solution with the fewest actions. Dijkstra’s algorithm finds minimum-cost paths when the costs of transitions are nonnegative. A* search uses both the cost already incurred and an estimate of the remaining cost to guide its exploration.
A* illustrates how prior knowledge can make planning more efficient. Instead of treating every unexplored possibility equally, it prioritizes paths that appear promising based on an estimate of how far they are from the goal. If the estimate satisfies the algorithm’s required conditions, A* can find an optimal path while avoiding some unnecessary exploration.
Other methods are better suited to large or complicated problems. Heuristic planners use practical rules to guide the search. Hierarchical planners divide a broad task into smaller subtasks. Optimization methods search for good combinations of decisions under specified constraints.
No search method eliminates complexity altogether. Some planning problems have so many possibilities that exhaustive search is impractical. Real systems therefore balance solution quality, computational cost, and the time available to make a decision.
How autonomous systems make decisions under uncertainty
A planner can produce a reliable sequence of actions only to the extent that it can predict the results of those actions. In real environments, information is incomplete, measurements are noisy, and outcomes are sometimes unpredictable.
A robot’s camera may fail to detect an object clearly. A self-driving vehicle may be unable to determine whether a pedestrian intends to cross. A software agent may not know whether a requested operation will succeed until it tries.
These situations require reasoning under uncertainty.
One important distinction is between a fully observable environment and a partially observable one. In a fully observable environment, the system has access to the information needed to identify the relevant state. In a partially observable environment, it cannot directly determine everything that matters.
When the state is uncertain, the system may maintain a belief state: a representation of what it currently considers possible and how likely different possibilities are. A robot searching for a misplaced object, for instance, might assign different probabilities to several locations based on its observations.
As new information arrives, the system updates its belief state. This process allows it to choose actions based not only on what it knows, but also on what it does not know.
A decision can be valuable because it gathers information, even when it does not immediately advance the main goal. A robot might move to a better viewing position before attempting to grasp an object. An autonomous vehicle might slow down to improve its ability to interpret an uncertain situation. A digital agent might inspect a file before deciding how to modify it.
This is an example of active information gathering. The system chooses an action partly for the information that action is expected to provide.
Expected utility and risk-sensitive decisions
When several outcomes are possible, a decision-maker can compare them using expected value or expected utility. Expected utility is calculated by considering the value of each possible outcome and weighting that value by the probability of the outcome occurring.
For example, a route may have a high probability of being fast but a small probability of causing a substantial delay. Another route may be slower on average but more predictable. Which route is preferable depends on the probabilities, the consequences of delays, and the system’s priorities.
Expected utility is useful, but it does not capture every consideration automatically. A system may need explicit safety constraints, limits on acceptable losses, or a preference for reliable outcomes when the cost of failure is unusually high.
In a safety-critical application, a low-probability hazard cannot necessarily be justified by a favorable average outcome. The system may be required to reject an action whenever it violates a safety constraint, regardless of its expected benefit.
Uncertainty also differs from variability that is well understood. A system may have a reliable probability model for a familiar situation, yet face a new environment in which that model is inaccurate. In the second case, calculating expected utility precisely does not make the decision reliable. The underlying estimates may be wrong.
The main approaches to AI planning and decision-making
There is no single method used by every autonomous system. Different approaches work well for different combinations of task complexity, uncertainty, available data, and safety requirements.
Classical planning represents actions through explicit preconditions and effects. A precondition is something that must be true before an action can occur; an effect is a change that the action produces. For example, a robot might be able to pick up an object only when it is close enough and its gripper is open. Once the object is grasped, the object’s location and the gripper’s state change.
This approach is especially useful when actions and constraints can be described clearly. It makes plans relatively easy to inspect and allows a system to reason systematically about whether a sequence of actions is feasible. Its limitations become more apparent when the environment is unpredictable or the action model is incomplete.
Rule-based decision systems use explicit conditions to select actions. A rule might instruct a machine to stop when a sensor detects an obstacle within a specified distance. Rules are understandable and can be useful for enforcing clear requirements, but large collections of rules can become difficult to coordinate. They may also struggle with situations that were not anticipated when the rules were written.
Optimization-based planning treats action selection as a mathematical problem. It seeks a solution that minimizes or maximizes a specified objective while respecting constraints. This is useful for scheduling, logistics, energy management, and motion planning. Its effectiveness depends on the quality of the mathematical model and whether the problem can be solved within available computational resources.
Machine learning offers another route. Rather than relying entirely on hand-written rules or a complete model of the environment, a system can learn patterns from data and use them to predict outcomes or choose actions.
These approaches are not mutually exclusive. An autonomous system might use a learned model to predict how an environment will respond, an optimizer to choose a route, and explicit safety rules to prevent prohibited actions.
How reinforcement learning teaches systems to choose actions
One influential approach to decision-making is reinforcement learning, in which an agent learns through interactions with an environment.
The agent observes a state, takes an action, and receives feedback in the form of a reward or penalty. Over repeated interactions, it attempts to learn a strategy that produces higher cumulative reward. The reward may represent immediate success, but it can also account for outcomes that emerge only after many steps.
The learned strategy is called a policy. A policy specifies which action to choose in a given state, or how to assign probabilities to available actions. A simple policy might tell a robot to move toward a target when the path is clear and stop when a hazard is detected. A learned policy can represent much more complicated behavior.
Reinforcement learning must often address the difference between short-term and long-term benefit. An action that produces little immediate reward may lead to a better result later. A system that learns to navigate, for example, may need to accept a temporary detour to avoid a situation that would otherwise prevent it from reaching its destination.
This creates a challenge known as the exploration-exploitation trade-off. Exploration involves trying actions to learn about their consequences. Exploitation involves choosing actions that already appear promising based on current knowledge. Too little exploration can leave the system with a poor strategy; too much can waste resources or introduce unacceptable risk.
Training often takes place in simulation or another controlled environment because unrestricted experimentation in the real world may be expensive or dangerous. However, a policy that performs well in simulation may fail when deployed in a physical environment with different lighting, friction, sensor characteristics, or unexpected obstacles. This gap between training conditions and real-world conditions is one of the important challenges in applying learned decision systems.
Reinforcement learning also depends heavily on reward design. If the reward does not accurately reflect the intended objective, the agent may discover a way to obtain high reward while behaving in an undesirable manner. A system rewarded for completing tasks quickly, for instance, may neglect aspects of quality or safety unless those requirements are represented adequately.
For these reasons, reinforcement learning is often combined with other methods rather than treated as a complete solution to autonomous decision-making.
How planning and learning work together
Planning and learning contribute different strengths. Planning can reason explicitly about action sequences and constraints, while learning can help systems recognize patterns, estimate unknown dynamics, and improve decisions from experience.
A robot may learn how its wheels respond to different surfaces, then use a planner to select a route that accounts for those predictions. A digital agent may use a learned model to estimate whether an operation will succeed, then construct a sequence of tool calls to complete a task.
In model-based reinforcement learning, an agent learns or uses a model of how actions change the environment and uses that model to guide decisions. In model-free reinforcement learning, it learns a policy or value estimates without necessarily constructing an explicit model of the environment’s transitions.
A value function estimates how desirable a state is, or how much future reward can be expected from a state or state-action pair. Such estimates can help a system compare options without searching through every possible future in detail.
Hybrid systems can also separate high-level and low-level decisions. A high-level planner may decide that a robot should pick up a particular object and deliver it to a shelf. A lower-level controller determines how to move its joints or wheels to carry out each step. The planner handles the task structure, while the controller manages the detailed physical behavior.
This division is important because selecting a sensible action sequence is not the same as executing it successfully. A robot may plan a collision-free path but fail to follow it because of wheel slip or inaccurate motion estimates. A controller may compensate for small deviations, while a planner may need to generate a new path if the deviation becomes substantial.
The result is a layered architecture in which goal selection, planning, control, and learning support one another.
Why autonomous systems must replan
A plan is based on assumptions about the current state and the consequences of future actions. When those assumptions stop matching reality, the plan may no longer be useful.
Autonomous systems therefore often use receding-horizon planning, also called model predictive control in a common control-theory formulation. The system plans over a limited future horizon, executes only the first action or a short portion of the plan, observes the new state, and plans again.
This method combines forward-looking reasoning with frequent correction. Rather than committing to a long sequence that may become obsolete, the system repeatedly updates its decisions as new information becomes available.
Consider a vehicle traveling toward an intersection. Its initial plan may assume that traffic will continue moving normally. If a vehicle suddenly stops ahead, the autonomous system must respond to the changed conditions. It may brake, wait, or choose another maneuver, depending on the situation and applicable safety constraints.
Replanning can also be triggered by the failure of an action. A robot might discover that a door is locked, a package is unavailable, or a route is blocked. The original goal may still be achievable, but the system needs a different sequence of actions.
Replanning is not always the best response. In a rapidly changing environment, repeatedly solving a large optimization problem may be too slow. A system may instead use a precomputed fallback, a reactive control rule, or a simplified local planner. The appropriate balance depends on how quickly decisions must be made and how much computation is available.
How autonomous vehicles combine planning and control
Autonomous driving illustrates why real-world decision-making requires several methods to operate together.
A vehicle must estimate its position, identify road users and obstacles, interpret traffic signals, predict how other road users might move, and choose a safe path. It must then control steering, acceleration, and braking to follow that path.
These responsibilities are related but distinct. Perception estimates what is happening. Prediction estimates what may happen next. Planning chooses a course of action. Control translates that choice into physical commands.
The vehicle may generate several candidate trajectories, each representing a possible path through space over time. It can evaluate those trajectories against objectives and constraints such as collision avoidance, road boundaries, traffic rules, passenger comfort, and progress toward the destination.
Prediction introduces uncertainty because other road users are independent decision-makers. A pedestrian may cross, wait, or change direction. A neighboring vehicle may merge or remain in its lane. The system cannot assume that every other actor will follow a single predicted path.
A robust driving system must therefore account for plausible alternative outcomes, monitor the situation as it develops, and preserve the ability to respond safely. It may choose a conservative action when uncertainty is too great to support a more aggressive maneuver.
Low-level control then attempts to follow the selected trajectory. Feedback from sensors helps correct deviations between the intended and actual motion. If the road changes or the trajectory becomes unsafe, the planning process must update its choice.
This example also shows why autonomous behavior cannot be judged solely by whether an algorithm finds an optimal path in a simplified model. The model may omit important hazards, perception may be imperfect, and the physical vehicle may not respond exactly as expected. Safety depends on the performance of the entire system.
How AI decision-making is evaluated
An autonomous system should be evaluated not only on whether it reaches its goal, but also on how it behaves while trying to reach it.
Task success is an obvious measure, but other criteria may be equally important. A robot’s performance might be assessed by completion time, energy consumption, damage to objects, or the frequency of failed operations. A vehicle’s performance may depend on safety, rule compliance, reliability, and its ability to handle unusual situations.
Decision quality is often difficult to measure because the best action may not be known in advance. In a controlled environment, the system can be compared with an optimal solution or a trusted baseline. In a complex real-world environment, evaluators may instead measure outcomes across many scenarios and examine whether the system behaves consistently under changing conditions.
Testing should include situations that differ from those encountered during development. Otherwise, an apparently successful system may simply have learned to perform well under familiar conditions. For learned systems, it is particularly important to examine whether performance deteriorates when inputs, environments, or task requirements change.
A further concern is the distinction between average performance and rare failures. A system may perform well in most situations but still fail in an unusual circumstance with serious consequences. Evaluating autonomous decision-making therefore requires attention to failure modes, not just aggregate success rates.
For systems that affect people or operate expensive equipment, decision processes may also need to be auditable. Developers should be able to identify the objectives, constraints, assumptions, and relevant information that influenced an action. This does not mean every learned model can provide a complete explanation of its internal computations. It means the surrounding system should support meaningful investigation of its behavior wherever possible.
The limits of autonomous decision-making
AI planning and decision-making remain limited by the information, models, objectives, and computational resources available to a system.
An inaccurate model can make a poor action appear optimal. If a planner underestimates how long a task takes, ignores a relevant hazard, or assumes that an unavailable resource can be accessed, it may construct a plan that cannot succeed. More sophisticated search cannot fully compensate for incorrect assumptions about the world.
Incomplete objectives create a different problem. A system can optimize what it has been instructed to optimize, but that objective may not capture every human preference or practical requirement. A scheduling system that minimizes travel time may produce an inconvenient route if accessibility or fairness is not included in its criteria. A task-completion agent may make unnecessary changes if it is not constrained to preserve unrelated information.
These problems are particularly difficult when objectives conflict. Speed, safety, cost, reliability, and flexibility may point toward different actions. There may be no single choice that maximizes every desirable property, so system designers must decide which requirements are mandatory and how remaining trade-offs should be handled.
Computational limits also matter. A theoretically optimal plan may take too long to calculate to be useful. In time-sensitive environments, a sufficiently good decision delivered quickly can be more valuable than an optimal decision delivered after the opportunity has passed.
Finally, autonomous systems can encounter situations outside the range of their training or design assumptions. A system may recognize familiar patterns without understanding the broader context in the way a person does. It may also be unable to determine when its own model is unreliable unless that capability has been explicitly developed and tested.
For these reasons, autonomy does not eliminate the need for human judgment. Human involvement may be necessary to define objectives, establish safety boundaries, monitor performance, review failures, and intervene when a system encounters conditions it cannot handle reliably.
AI planning and decision-making provide a technical foundation for autonomous behavior, but their success depends on more than selecting an action that looks best according to a mathematical objective. The system must represent the situation accurately enough, reason about relevant consequences, account for uncertainty, execute its choices reliably, and recognize when circumstances require a different plan.