This website uses cookies

Read our Privacy policy and Terms of use for more information.

Updated September 11, 2026. The Mamba architecture, introduced in December 2023, has grown from an intriguing alternative to attention into a broader family of sequence models. We briefly discussed the original question in our weekly digest: “What is Mamba and can it beat Transformers?” The answer is more interesting now. Mamba did not replace the Transformer, but it changed how researchers think about efficient sequence modeling – and its ideas increasingly appear in hybrid systems.

This article, part of our AI 101 series, follows Mamba from its roots in state space models to Mamba-2 and Mamba-3. It explains what selective state spaces solve, where the comparison with Transformers is fair, and where it is not.

In today’s episode, we will cover:

  • Sequence modeling, Transformers, and foundation models

  • Transformer Limitations: Quadratic Scaling and Context Windows

  • State Space Models (SSMs): An Alternative to Transformers

  • Mamba vs Transformer: key differences

  • Where Mamba is used today

  • Mamba-2 and Mamba-3: what changed

  • FAQ and further reading

Sequence Modeling: From RNNs to Transformers

First, let's understand the context of foundation models, transformers, and related concepts. It’s called sequence modeling, an area of research driven by sequential data, meaning it has some inherent order inside. It can be text sentences, images, speech, audio, time series, genomics, etc.

In the 1990s and 2000s, sequence modeling advanced significantly due to the rise of neural network architectures and increased computing power. Models like Long Short-Term Memory (LSTM) and recurrent neural networks (RNNs) became the standards. These models are excellent at handling variable-length sequences and capturing long-term dependencies. However, their sequential data processing nature makes it challenging to parallelize operations within a single training example.

In 2017, the introduction of the transformer architecture marked a significant change in sequence data processing. Transformers use attention mechanisms, eliminating the need for the recurrence or convolution used in earlier models. This design enhances parallelization during training, making transformers more efficient and scalable for handling large datasets.

The self-attention mechanism in transformers computes the relevance of all other parts of the input sequence to a particular part. This is done for all parts of the sequence simultaneously, which is inherently parallelizable. Each attention head can independently calculate the attention scores for all positions in the input sequence, making these operations run concurrently. This differs drastically from RNNs where outputs are computed one after the other. This results in markedly faster training times and the ability to handle larger datasets more efficiently.

Okay, so Transformers really became the backbone of modern foundation models. As ChatGPT would say, they really propelled their development! But transformers are not the only ones in sequence modeling, and it turned out they have some problems with long context. So it’s time to turn to our main topic – Mamba.

Transformer Limitations: Quadratic Scaling and Context Windows

While Transformers have many advantages, they are not without drawbacks. Standard full self-attention lets every token interact directly with every other token in the available context. That direct access is powerful, but its training-time compute and memory grow quadratically with sequence length. During autoregressive generation, the key-value cache also grows as the context grows. A context window is therefore a design and deployment limit – not a rule that attention can see only nearby tokens. One line of work attacks this cost inside attention itself: Linformer (Wang et al., 2020) projects keys and values to a lower-rank space, so inference time stays roughly flat as the sequence grows.

Linformer vs Transformer inference time: low-rank key and value projection keeps latency flat as sequence length grows

Image Credit: Stanford University course "Natural Language Processing
with Deep Learning CS224N/Ling284

Moreover, increasing the size of this context window exacerbates computational demands, leading to quadratic scaling with respect to the window length. Specifically, if the context length is x, the computational resources required to scale is x2.

That’s a lot.

State Space Models (SSMs): An Alternative to Transformers

Recently, state space sequence models (SSMs), including a variant known as Mamba, have emerged as potential successors to transformers in sequence modeling. These models address some of the fundamental limitations inherent to transformers.

It's important to recognize that SSMs like Mamba specifically tackle certain constraints of transformers, yet many other approaches also seek to enhance transformer efficiency. The paper "Efficient Transformers: A Survey" extensively discusses these methodologies and illustrates them in the image below. For example, last time, we talked about Mixture-of-Experts (MoE), categorized under "Sparse" models in the provided Venn diagram.

State space models (SSMs)

Overall, the approach of using state space models draws us back to dynamic systems which are described using differential equations. A dynamic system is any system, man-made, physical, or biological, that changes in time. Think of the Space Shuttle in orbit around the Earth, an ecosystem with competing species, the nervous system of a simple organism, or the expanding universe, these are all systems.

State space models are extensively used in many other sciences but are relatively overlooked in machine learning. They struggled even on simple tasks before introducing special state matrices*. With these matrices, state space models achieve exceptional performance.

*State matrices are fundamental components of state-space models. These matrices form part of a mathematical framework that describes how the state of a system evolves over time in response to inputs.

Structured state space sequence models were proposed in 2021 by three researchers from Stanford University: Albert Gu, Karan Goel, and Christopher Ré, and were an extension of their previous work. These models blend elements from recurrent neural networks (RNNs) and convolutional neural networks (CNNs), drawing from classical state space principles. Nonetheless, they tend to underperform with discrete, dense data like text.

What Is Mamba? Selective SSM Architecture Explained

In response to early SSMs, researchers from Carnegie Mellon University and Princeton introduced the Mamba model, an innovative selective state space model that overcomes the drawbacks of both Transformers and traditional SSMs.

The Mamba architecture incorporates selective state space models into a neural network framework to eliminate traditional components such as attention and Multilayer Perceptron (MLP) blocks.

Mamba Architecture: Selective SSM, Hardware-Aware Algorithm, and Mamba Block

  • Selective SSMs: Similar to structured SSMs, selective SSMs in Mamba are standalone sequence transformations integrated into neural networks. These SSMs allow parameters to adapt based on input, optimizing information propagation throughout the sequence. This feature supports Mamba's capacity for high-efficiency content-based reasoning across diverse modalities such as language, audio, and genomics.

  • Simplified structure: Mamba simplifies traditional complexity by merging the linear attention and MLP blocks into the "Mamba block." The architecture simplifies further by using several identical Mamba blocks stacked one after another instead of mixing different types of blocks. This uniformity contributes to the simplicity and efficiency of the model. See the diagram above.

  • Hardware-aware algorithm: To tackle the computational inefficiencies caused by SSMs, researchers came up with a clever hardware-aware parallel algorithm. Instead of using convolution, this algorithm uses a scan operation, which helps in managing state expansion within the memory hierarchy of modern GPUs. This boosts speed and cuts down on memory overhead.

Mamba vs Transformer: Key Differences

Dimension

Mamba-style SSM

Transformer

Sequence computation

Linear-time recurrence or scan

Full attention is quadratic in sequence length

State during decoding

Fixed-size recurrent state

KV cache grows with the context

Access to earlier information

History is compressed into the state

Tokens can attend directly to earlier tokens

Practical maturity

Performance depends heavily on specialized kernels

Mature tooling and highly optimized hardware support

Where it is attractive

Long or streaming sequences and memory-constrained decoding

General-purpose quality, exact retrieval, and ecosystem support

The short answer is efficiency over long sequences – but that is not the same as saying Mamba is always better. Full attention compares tokens directly; Mamba compresses what came before into a recurrent state. That gives Mamba linear sequence scaling and a fixed-size decode state, but it can make exact recall and state tracking harder. The practical choice depends on the task, the hardware, and the implementation. Increasingly, model builders combine attention with Mamba-style layers instead of treating the two as mutually exclusive.

In the original paper’s experiments, the 3B-parameter Mamba model outperformed Transformers of the same size and matched models reported as roughly twice its size on the evaluated language-modeling setup. That was an important result, but it should not be read as a universal ranking against today’s Transformer models.

Its main advantages:

  • Linear sequence scaling: Mamba processes a sequence through a recurrence or parallel scan whose work grows linearly with sequence length. Standard full attention grows quadratically during training.

  • Compact decoding state: Autoregressive generation carries a fixed-size recurrent state rather than a KV cache that grows with every generated token.

  • Hardware-aware execution: The original implementation introduced a selective-scan algorithm designed for GPUs. Its paper reported up to five times higher inference throughput in its tested setup; real gains vary by model size, kernel, hardware, batch size, and sequence length.

  • Selective state updates: Unlike earlier linear time-invariant SSMs, Mamba lets the input control which information enters, remains in, or leaves the state. That selectivity is central to its performance on content-rich sequences. The weights that govern the state, however, stay fixed once deployed. A different line of work, test-time training, updates those parameters on the data the model meets while running.

The advent of the Mamba architecture is a fascinating update to transformers integrating state space models into sequence modeling. It’s a promising approach that could potentially allow us to achieve new levels of efficiency and scalability in foundation models. And it demonstrates again: the innovation is based on previous research and it’s important to know your ML history.

Where Is Mamba Used Today?

Mamba’s clearest impact is not a single pure-Mamba winner. It is the spread of state-space layers across language, audio, genomics, vision, video, medical data, and robotics – and especially inside hybrid models. AI21’s Jamba family mixes Mamba and attention layers. TII’s Falcon-H1 and NVIDIA’s Nemotron-H family follow the same broad idea: use recurrent state-space layers for efficient sequence processing, then keep attention where direct token access helps. Hybrids are not the only direction. Mixture-of-Mamba keeps the blockitself but gives each modality its own projection parameters instead of wrapping attention layers around it.

That is the useful 2026 conclusion. Mamba is not “the new Transformer.” It is now one of the main architectural tools available to model builders, and hybrid designs show how the two approaches can complement each other.

Mamba-2 and Mamba-3: What Changed

Mamba-2, published in 2024, introduced State Space Duality: a framework that connects structured state space models with forms of attention. The resulting SSD layer was designed to train more efficiently and support larger state sizes. Its authors reported a core layer two to eight times faster than Mamba’s while remaining competitive on their small- and medium-scale language-modeling tests.

Mamba-3, released in March 2026, starts from an inference-first question: can a linear model become more expressive without giving up efficient decoding? It combines a more expressive recurrence, complex-valued state updates, and a multi-input, multi-output formulation. At the 1.5B scale, the paper reports better retrieval, state tracking, and downstream accuracy than the linear-model baselines it tested. These are promising research results, not evidence that Mamba-3 has displaced frontier Transformers.

Thank you for reading!

FAQ

What is Mamba in AI?

Mamba is a selective state space architecture for sequence modeling. It uses input-dependent recurrent dynamics instead of full self-attention in its core block, allowing sequence computation to scale linearly with length.

Is Mamba better than a Transformer?

Not universally. Mamba has attractive memory and scaling properties for long or streaming sequences. Transformers retain direct token-to-token access, mature tooling, and strong general-purpose performance. The best choice depends on the workload, and many current systems use both.

What is the difference between Mamba, Mamba-2, and Mamba-3?

Mamba introduced selective state spaces. Mamba-2 connected structured SSMs and attention through State Space Duality and improved training efficiency. Mamba-3 added a more expressive recurrence, complex-valued state updates, and a MIMO formulation designed around efficient inference.

Does Mamba use attention?

The core Mamba block does not require standard full self-attention. Hybrid architectures can – and often do – combine Mamba-style layers with attention.

What are Mamba models used for?

Mamba variants are studied and deployed in language models, speech and audio, genomics, vision, video, medical imaging, robotics, and other long or streaming sequence tasks.

Bonus: Mamba-related research papers worth reading:

Bonus

Reply

Avatar

or to participate

Keep Reading

View more
caret-right