Exploring the breakthrough innovations shaping our world. From AI infrastructure and robotics to biotech, quantum computing, and spatial tech.
When engineers evaluate an AI accelerator, two numbers often attract the most attention: compute performance and memory bandwidth. Compute performance describes how many mathematical operations a processor can perform, while memory bandwidth describes how quickly data can move between the accelerator and its high-bandwidth memory system.
The first number is easier to market. A new accelerator can promise dramatically more floating-point or tensor operations than its predecessor, and those figures make for impressive headlines. But raw compute is only useful when the processor can be supplied with enough data to keep its execution units busy.
That is where the memory wall enters the picture. Modern AI accelerators can perform enormous numbers of calculations, yet some workloads remain limited by how quickly data can move through the memory hierarchy. High-Bandwidth Memory, or HBM, has become an important part of the solution because it provides extremely wide memory interfaces and very high bandwidth in a compact package.
HBM does not eliminate the memory wall. It moves the boundary.
The memory wall describes a long-standing problem in computer architecture: compute capability can improve much faster than the latency and bandwidth characteristics of external memory systems.
An easy way to picture the problem is to imagine a professional kitchen. The chefs can prepare ingredients extremely quickly, but the ingredients are stored in a distant pantry. If deliveries cannot keep pace with the kitchen, the chefs eventually spend time waiting rather than cooking.
The same basic problem can appear inside an AI accelerator. The execution units may be capable of performing enormous numbers of mathematical operations, but those operations still require operands, weights, activations, and intermediate data to be available at the right time.
This is why performance cannot be judged solely by FLOPS or tensor-operation figures.
The Roofline model provides one useful way to think about this trade-off. It relates achievable performance to arithmetic intensity—the amount of computation performed for each byte of data moved. A workload with relatively low arithmetic intensity can become bandwidth-bound, meaning that adding more compute capability does little unless the memory system can also deliver more data.
That distinction is especially important in AI because different neural-network operations can have very different compute-to-memory characteristics.

Traditional computers use several kinds of memory, including DDR system memory and GDDR memory in graphics systems. These technologies remain extremely useful and have not become obsolete simply because AI accelerators use HBM.
The problem is scale.
A modern accelerator may contain thousands of parallel execution units that can consume data at extraordinary rates. Supplying those units through conventional board-level memory interfaces becomes increasingly difficult as bandwidth requirements rise.
Several physical constraints contribute to the problem.
One way to increase memory bandwidth is to increase the width of the memory interface. Another is to increase signaling speed. Both approaches have practical limits.
Adding more interface pins consumes package and board space, while pushing signaling rates higher can increase power consumption and create signal-integrity challenges.
HBM takes a different approach: rather than relying primarily on very high signaling rates, it uses an exceptionally wide interface.
Conventional memory can be physically separated from the processor by relatively long electrical connections on a circuit board. HBM is positioned much closer to the accelerator through advanced packaging.
Shorter connections and a much wider interface make it possible to move large amounts of data without relying entirely on extremely high signaling frequencies.
Moving data is not free. As AI systems scale, the energy required to move information through the memory hierarchy becomes an important part of overall system efficiency.
HBM's architecture is designed to deliver high bandwidth while keeping the distance between the memory stacks and accelerator relatively short. That makes it attractive for high-performance computing and AI workloads.

HBM changes the physical relationship between DRAM and the processor.
Instead of treating memory as a collection of conventional modules connected through relatively narrow interfaces, HBM uses vertically stacked DRAM dies and a very wide interface to the processor package.
Three technologies are particularly important.
HBM stacks multiple DRAM dies vertically. Rather than spreading memory chips across a large area, the architecture builds memory capacity upward.
This approach improves memory density while creating a structure that can communicate through many parallel connections.
The stacked dies communicate through through-silicon vias, or TSVs. These microscopic vertical connections allow electrical signals to travel between different layers of the memory stack.
TSVs are one of the key technologies that make high-density stacked memory possible.
The other major advantage is interface width.
HBM uses a much wider interface than conventional memory technologies. Instead of depending solely on higher clock speeds to increase bandwidth, it moves more data in parallel.
That combination—stacked dies, short connections, and a very wide interface—is what allows modern HBM systems to reach bandwidth levels measured in terabytes per second.
The numbers are not merely theoretical. For example, NVIDIA's H200 accelerator uses HBM3e and is specified with 141 GB of HBM3e memory and 4.8 TB/s of memory bandwidth.
That illustrates why HBM has become so closely associated with modern AI accelerators.

HBM dramatically increases available bandwidth, but higher bandwidth does not make memory traffic disappear.
AI workloads continue to grow in several directions at once. Models can become larger, context windows can become longer, batch sizes can change, and serving systems may need to handle many requests concurrently.
The result is a continuing demand for efficient movement of data.
Consider autoregressive language-model inference. The model's weights need to remain accessible throughout token generation, while activations and other intermediate data move through the accelerator's memory hierarchy.
For some inference workloads, particularly those with relatively low arithmetic intensity, moving weight data can become a major factor in performance.
This is one reason memory capacity and bandwidth matter alongside raw compute.
It also explains why quantization can have such a meaningful effect. Representing weights with fewer bits reduces their memory footprint and can reduce the amount of data that needs to move through the memory hierarchy.
Training is more complicated.
The system must handle weights, activations, gradients, optimizer states, and communication between accelerators. Some operations are highly compute-intensive, while others can become limited by memory traffic.
The bottleneck can therefore shift from one part of the system to another depending on the model architecture, batch size, sequence length, precision, and implementation.
That is why it is more accurate to say that HBM bandwidth is a critical performance resource rather than the universal bottleneck for every AI workload.

Hardware is only part of the answer.
If an algorithm repeatedly moves the same data between HBM and faster on-chip memory, simply adding more compute units may not solve the problem. Reducing unnecessary data movement can sometimes produce a larger practical improvement.
FlashAttention is a good example.
The FlashAttention work introduced an IO-aware approach to attention that reduces reads and writes between GPU HBM and on-chip SRAM by organizing computation into tiles. The underlying idea is straightforward: if memory traffic is limiting performance, changing how the computation accesses memory can be just as important as increasing raw compute capacity.
This is an important shift in how AI performance is understood.
The question is not simply:
How many operations can the accelerator perform?
It is also:
How much useful work can the accelerator perform for every byte that moves through the memory hierarchy?
Quantization attacks the same problem from another direction.
A model represented with 16-bit values requires more memory and memory traffic than the same model represented with 8-bit or 4-bit values. Lower precision can therefore make a model smaller and reduce the amount of data that must be transferred.
The trade-off is accuracy and numerical behavior.
Not every model or workload responds identically to lower precision, so quantization is not a universal replacement for additional memory bandwidth. But it demonstrates an important principle: improving AI performance does not always mean adding more silicon.
Sometimes the better solution is to move less data.
There is another reason HBM has become strategically important: manufacturing it is considerably more complex than producing a conventional DRAM module.
The process involves stacked memory dies, microscopic vertical connections, advanced packaging, and demanding thermal and electrical requirements.
Higher stack counts can increase capacity, but they also make manufacturing and thermal management more challenging.
The packaging layer matters just as much as the memory dies themselves because the accelerator and its memory must work together at very high bandwidth.
This creates a supply-chain constraint that is different from simply manufacturing more conventional DRAM.
Increasing HBM capacity and bandwidth is one path forward, but it is not the only one.
Future AI systems will likely combine several approaches:
more efficient memory hierarchies;
better on-chip caches and SRAM usage;
lower-precision computation;
algorithms that minimize unnecessary data movement;
faster accelerator-to-accelerator interconnects;
more advanced packaging;
and potentially optical technologies for selected high-bandwidth links.
Co-packaged optics is particularly interesting for interconnects because electrical links become increasingly difficult to scale as bandwidth and distance increase. Optical approaches may help with bandwidth density and energy efficiency, but they do not eliminate latency or system-level engineering constraints.
The broader trend is clear: improving AI performance increasingly requires attention to data movement, not just arithmetic throughput.

The rise of HBM has changed the balance between compute and memory.
Modern accelerators can access far more memory bandwidth than earlier generations, allowing them to keep increasingly powerful compute engines supplied with data. But every improvement in compute creates pressure for another improvement somewhere else in the memory hierarchy.
That is why the memory wall remains relevant.
It is not a single physical barrier, and HBM is not a magic solution. Some workloads are compute-bound. Others are bandwidth-bound. Still others are limited by communication between accelerators, memory capacity, synchronization, or software efficiency.
The important lesson is that AI performance depends on the entire path that data takes through the system.
A faster processor helps only when the rest of the architecture can keep it supplied. HBM addresses one of the most demanding parts of that equation by bringing extremely wide, high-bandwidth memory close to the accelerator.
As AI systems continue to scale, the competition will not simply be about building more powerful processors. It will also be about moving the right data to the right place, at the right time, with as little wasted movement as possible.