Recurrent neural networks (RNNs) are a type of artificial neural network designed to process information that unfolds over time or arrives in a particular order. Unlike models that treat each input as an independent item, RNNs can use information from earlier steps to interpret what comes next. This makes them well suited to tasks involving language, speech, time-series measurements, and other sequential data.
The central idea is that an RNN maintains a changing internal state as it processes a sequence. At each step, it combines the current input with information carried forward from previous steps, producing an updated representation of what it has encountered so far. That state gives the network a limited form of memory, allowing the meaning of a new input to depend on the context established by earlier ones.
RNNs helped establish important principles for sequence-based artificial intelligence. Although newer architectures, particularly transformers, have become prominent in many language-processing applications, RNNs remain valuable for understanding how neural networks learn temporal patterns, how information can be represented across a sequence, and why remembering the past is a difficult computational problem.
Why sequential data requires a different approach
Many forms of data have an inherent order. The meaning of a sentence depends partly on the arrangement of its words. A spoken command develops through a succession of sounds. A patient’s heart-rate measurements form a time series, and the significance of a reading may depend on what happened before it. In each case, examining individual elements without considering their context can discard important information.
A conventional feedforward neural network processes information through a sequence of computational layers, moving from input to output without maintaining an internal state across separate inputs. It can learn relationships within a fixed-size input, including relationships among several words or measurements supplied together. However, it does not inherently carry information from one processing step to the next when those steps are presented separately.
An RNN addresses this limitation by introducing a recurrent connection. The network uses its previous hidden state—the internal representation of information accumulated so far—alongside the current input to calculate its next hidden state. The process repeats for each element in the sequence.
Consider the sentence, “The package arrived late because the delivery truck broke down.” Understanding the final explanation requires processing the words in order. An RNN can update its internal representation as each word arrives, allowing later words to be interpreted in light of earlier ones. The network does not necessarily preserve a perfect record of every word, but it can learn which aspects of previous inputs are useful for predicting or interpreting what follows.
This capacity makes RNNs natural candidates for sequential problems. Their distinctive feature is not simply that they process data one item at a time, but that they reuse an internal state to connect successive items.
How a recurrent neural network works
An RNN processes a sequence through repeated applications of the same basic computation. At each time step, it receives an input, combines that input with its previous hidden state, and calculates a new hidden state. It may then use that state to produce an output.
For example, when processing a sentence, the input at each step might be a numerical representation of one word. The hidden state summarizes information the network has retained from earlier words. As the sentence progresses, this representation changes to reflect both the new word and the context already accumulated.
The underlying computation is often expressed as:
ht=f(Wxxt+Whht−1+b)h_t = f(W_xx_t + W_hh_{t-1} + b)ht=f(Wxxt+Whht−1+b)
Here, xtx_txt represents the current input, ht−1h_{t-1}ht−1 is the previous hidden state, and hth_tht is the updated hidden state. The matrices WxW_xWx and WhW_hWh contain learned weights that determine how the current input and previous state influence the update. The term bbb is a bias, and fff is a nonlinear activation function that allows the network to learn more complex relationships than a simple linear transformation could represent.
These weights are reused at every time step. An RNN processing the tenth word of a sentence applies the same learned transformation used for the first word, rather than requiring a completely separate set of parameters for every position. This shared structure allows the model to handle sequences of different lengths, provided its input and output design supports them.
The hidden state acts as the network’s working memory. It is not a literal record of the sequence, nor does it necessarily correspond to a human-readable summary. Instead, training encourages the network to encode information that helps it perform its assigned task.
Depending on the application, the network can produce an output at every time step or only after processing the entire sequence. A language model, for instance, may use the current state to predict the next word or token. A sequence classifier may process a complete passage and use its final state to assign a category, such as identifying whether a message expresses a complaint.
Some RNNs also maintain a separate output representation at each step. The distinction between hidden states and outputs matters: the hidden state supports subsequent computation, while the output is the information the model exposes for its task.
How RNNs learn from sequential data
RNNs typically learn through a process called supervised learning, in which the network’s predictions are compared with target answers. Training adjusts the model’s weights to reduce a loss function, a mathematical measure of how far its predictions are from the desired results.
Suppose an RNN is trained to predict the next word in a sentence. At each step, it receives the preceding context and predicts a distribution of possible next words. The training process compares that distribution with the actual next word and calculates a loss. Repeating this process over many examples helps the network learn statistical patterns in word order, grammar, and context.
The learning process requires more than calculating an error at the final output. The network must determine how its internal weights contributed to that error. Neural networks generally do this using backpropagation, which calculates how changes in model parameters would affect the loss. For recurrent networks, this process is extended across the sequence through a method called backpropagation through time.
To understand this method, imagine unfolding the recurrent computation into a chain of copies, one for each time step. Although the chain contains many computational stages, all of them share the same underlying weights. Backpropagation calculates how the error depends on those weights by tracing the effects of each stage through the preceding stages.
The resulting gradients—numerical measures that indicate how parameters influence the loss—are used by an optimization algorithm to update the weights. Over repeated training cycles, the model becomes better at capturing patterns that support its objective.
This process creates a fundamental connection between memory and learning. An earlier input can affect a later prediction through the hidden states connecting the two steps. Training must therefore assign credit or blame to computations that may have occurred many steps before the observed error.
That requirement also introduces one of the main challenges of RNNs: learning relationships across long sequences can be difficult, even when the relevant information is present in the input.
Why ordinary RNNs struggle with long-term dependencies
A major limitation of traditional RNNs is their difficulty retaining useful information over long stretches of a sequence. This problem is known as the long-term dependency problem.
Consider a sentence that begins by introducing a person or object and refers to it again much later. Correctly interpreting the later reference may require information from the beginning. In time-series forecasting, a measurement might depend on a pattern that developed many steps earlier. A basic RNN can theoretically carry such information forward, but its learning process does not always make that practical.
The difficulty becomes particularly clear during backpropagation through time. To learn how an early input affects a much later prediction, the network must propagate gradient information backward through many recurrent steps. During this process, gradients can become extremely small or extremely large.
When gradients become very small, a problem called the vanishing gradient problem, the learning signal reaching earlier steps may be too weak to make meaningful adjustments. The network then struggles to learn how distant events influence later outcomes. When gradients grow excessively, the exploding gradient problem, training updates can become unstable and produce extreme changes in the model’s parameters.
These problems arise partly because the same recurrent transformations are applied repeatedly. The effect of a small change in the hidden state can shrink or expand at each step, depending on the learned weights and the network’s activation functions. Across a long sequence, those repeated effects can make optimization difficult.
An important distinction is that an RNN’s theoretical ability to carry information is not the same as its ability to learn to preserve that information reliably. A basic RNN can represent certain long-range relationships in principle, yet ordinary training may fail to discover or maintain them.
Researchers have developed several approaches to address these limitations, including improved recurrent architectures, gradient clipping to control exploding gradients, careful parameter initialization, and alternative sequence models. Among the most influential solutions are long short-term memory networks and gated recurrent units.
How LSTMs and GRUs improve recurrent memory
Long short-term memory networks, usually called LSTMs, were developed to make learning long-range dependencies more manageable. They retain the recurrent structure of an RNN but introduce a more elaborate memory mechanism that controls how information is stored, updated, and exposed.
An LSTM maintains a cell state in addition to its hidden state. The cell state provides a pathway through which information can be carried forward with relatively controlled modifications. The network uses learned gates to regulate this process.
A gate is a mechanism that determines how much information should pass through or be modified. The forget gate controls how much of the previous cell state to retain. The input gate regulates how much new information to incorporate. The output gate controls how much of the cell state contributes to the hidden state exposed to subsequent computations.
These gates are calculated from the current input and the previous hidden state. Their values vary with the data, allowing the network to preserve information in some circumstances and update it in others. For example, an LSTM may learn to retain a relevant feature across several steps while discarding details that are less useful for its prediction.
The advantage is not that an LSTM remembers everything indefinitely. Its memory is still finite, learned, and imperfect. Rather, its architecture provides more direct control over information flow, making it easier to preserve useful signals and learn dependencies over longer intervals than in many basic RNNs.
Gated recurrent units, or GRUs, offer a related approach with a somewhat simpler design. They use update and reset gates to regulate how new inputs interact with the existing state. Unlike a standard LSTM, a GRU does not maintain a separate cell state. Its more compact structure can reduce the number of parameters required, although whether it performs better than an LSTM depends on the task and training conditions.
Neither architecture is universally superior. Performance depends on the sequence length, the amount and quality of training data, the computational budget, the model’s size, and the nature of the patterns it must learn. LSTMs and GRUs remain useful examples of how architectural changes can make recurrent memory more effective.
How RNNs are used in language, speech, and time-series analysis
RNNs are particularly suited to tasks in which the order of observations carries meaningful information. Their applications span language processing, speech recognition, forecasting, and the analysis of measurements collected over time.
In natural language processing, RNNs have been used for language modeling, text classification, machine translation, and sequence labeling. Sequence labeling assigns a label to each element in a sequence, such as identifying which words in a sentence refer to people, organizations, or locations. Because the meaning of a word often depends on its surroundings, a recurrent hidden state can provide useful context.
In speech processing, audio is represented as a succession of measurements or extracted features. An RNN can use earlier acoustic information to help interpret the current segment, supporting tasks such as speech recognition. Speech is continuous, and the evidence for a particular sound or word may develop across multiple time steps. Recurrent processing offers a way to incorporate that evolving evidence.
Time-series analysis is another natural application. A time series is a sequence of observations indexed by time, such as daily electricity demand, hourly temperature readings, or successive measurements from an industrial sensor. An RNN can learn relationships between previous observations and later outcomes, making it useful for forecasting and anomaly detection.
In forecasting, for example, a network may receive a window of past measurements and predict one or more future values. The hidden state can encode patterns such as recurring fluctuations or changes in recent conditions. However, the network does not automatically understand the physical causes behind those patterns. It learns statistical relationships from the training data, which may fail to generalize when the underlying process changes.
RNNs have also been applied to event sequences, in which observations represent discrete events rather than regular measurements. Examples include user interactions, equipment failures, or successive actions in a process. Here, the intervals between events may matter as much as their order, so the model may need explicit information about elapsed time.
Across these applications, the same principle remains: previous observations can inform the interpretation of current ones. The architecture is useful when a task benefits from a learned representation of sequential context, but success depends on the suitability of the model, the data, and the training objective.
How RNNs compare with transformers
Transformers have become a major alternative to recurrent networks, particularly in natural language processing. Their defining mechanism, called attention, allows a model to determine which parts of an input are relevant to a particular computation.
In a conventional RNN, information passes through the sequence as the hidden state is updated step by step. A distant input can influence a later representation only through the recurrent computation connecting them, unless the architecture includes additional mechanisms. This makes the processing order intrinsic to the model’s design and can make training difficult over long sequences.
A transformer can instead calculate relationships between multiple positions in a sequence through attention. In self-attention, each position can use information from other positions that the model is permitted to access. This creates a more direct route for relating distant words or events, without requiring all information to pass through one chain of recurrent states.
Transformers also offer an important computational advantage during training: many sequence positions can be processed in parallel. A conventional RNN’s hidden state at one step depends on the previous step, so the recurrent computations themselves must follow their dependency order. This limits parallelization along the sequence, although other aspects of RNN training can still be parallelized.
Attention has its own costs. In standard full self-attention, computation and memory requirements can grow substantially with sequence length because the model evaluates relationships among many pairs of positions. RNNs, by contrast, can process a stream incrementally while maintaining a fixed-size hidden state, which can be useful when inputs arrive continuously or memory is constrained.
These differences do not mean that transformers always outperform RNNs on every sequential task. The most appropriate architecture depends on the application, the available resources, the length and structure of the data, and the importance of parallel processing or streaming operation. Other model families, including convolutional networks and state-space models, can also be effective for sequence problems.
The rise of transformers reflects their flexibility, strong performance across many language tasks, and suitability for large-scale training. RNNs nevertheless remain conceptually important because they illustrate a fundamental strategy for sequential computation: maintain an evolving state that allows earlier inputs to shape later outputs.
What an RNN’s memory can and cannot represent
The word memory can be misleading when applied to a neural network. An RNN does not remember in the same way a person recalls an experience. Its hidden state is a numerical representation learned through optimization, and the information it retains depends on the model’s architecture, parameters, training data, and objective.
Because the hidden state has a limited number of numerical components, the network must compress the history of the sequence into a representation of manageable size. It cannot generally preserve every detail of an arbitrarily long input without loss. Instead, training encourages it to retain information that improves performance on the task.
This compression can be useful. A model that predicts the next word may not need a verbatim copy of every preceding word. It may benefit more from retaining grammatical context, the topic of the passage, or information about a subject mentioned earlier. For a forecasting model, the most useful representation might encode recent trends or recurring patterns.
However, what the network preserves is not guaranteed to match what a human considers important. A model can overlook a subtle but consequential detail, rely on an incidental correlation, or perform well on familiar data while failing on unfamiliar examples. Its internal state is not necessarily interpretable, and an accurate prediction does not by itself establish that the model has learned the underlying causal process.
The model’s memory also depends on the task’s design. A sequence classifier that uses only the final hidden state may lose information that would have been useful at intermediate steps. A model that produces an output at every time step may preserve and use context differently. Bidirectional recurrent networks, which process a sequence in both forward and reverse directions, can incorporate information from both earlier and later positions when the full sequence is available. They are therefore useful in some offline language and labeling tasks, but they cannot be used in the same way for strictly real-time prediction when future inputs are not yet known.
These distinctions show why sequential modeling involves more than selecting an architecture. Researchers must decide what counts as an input step, which information the model can access, what output it should produce, and how success should be measured.
Why recurrent neural networks remain important
RNNs demonstrate how a neural network can use a shared computation to connect observations across time. Their recurrent structure provides a straightforward way to represent sequential context, while their training process reveals how difficult it can be to learn relationships that extend far into the past.
Their limitations have also shaped the development of modern AI. The challenges of vanishing gradients motivated architectures such as LSTMs and GRUs. The constraints of sequential computation helped clarify the benefits of architectures that process many positions in parallel. Together, these developments illustrate a broader principle in machine learning: a model’s architecture determines not only what it can represent, but also how easily it can learn useful representations from data.
RNNs are not simply an earlier version of today’s language models. They embody a distinct approach to sequence processing, one that can still be appropriate when inputs arrive continuously, compact recurrent state is desirable, or the computational demands of a task favor incremental processing.
Understanding RNNs also provides a foundation for understanding AI more broadly. Many intelligent systems must interpret information in context, distinguish relevant patterns from noise, and use earlier observations to guide later decisions. Recurrent neural networks address these challenges by turning a sequence into a series of linked computations, each informed by a learned representation of what came before.