Exploring the breakthrough innovations shaping our world. From AI infrastructure and robotics to biotech, quantum computing, and spatial tech.
For much of the modern computing era, general-purpose systems have followed a basic architectural principle: processing and main memory are separate parts of the system. A processor retrieves data from the memory hierarchy, performs calculations, and writes results back to storage or another level of memory. This arrangement provides enormous flexibility, but it also creates a persistent engineering challenge. As processors have become capable of performing more operations in parallel, supplying them with data quickly and efficiently has become increasingly important.
This challenge is often described as the memory wall. The problem is not simply that memory is "slow" or that processors are "fast." It involves the combined effects of memory latency, bandwidth, cache capacity, interconnects, and the energy required to move data between different parts of a computing system. In-memory computing and processing-in-memory (PIM) architectures approach the problem from a different direction: instead of repeatedly bringing large amounts of data to a separate processor, they place selected computational operations closer to where the data is stored. The result is not a replacement for conventional computing, but a specialized architectural strategy for workloads where data movement has become a significant constraint.

A modern processor rarely works with data in only one location. Information can move through registers, multiple levels of cache, main memory, accelerators, and storage, with each level offering a different combination of capacity, latency, bandwidth, and energy consumption. Caches reduce the need to access slower memory, while techniques such as prefetching and parallel memory channels help keep processors supplied with data. These mechanisms are highly effective, but they cannot eliminate the underlying cost of moving information through the system. When a workload repeatedly accesses large datasets with limited opportunities for reuse, memory behavior can become more important to performance than the arithmetic itself.
This is especially relevant to memory-bound workloads, where execution is constrained by data access rather than by the number of arithmetic operations available. Graph analytics, recommendation systems, database operations, some scientific workloads, and selected machine-learning tasks can exhibit this behavior. The exact balance varies considerably by algorithm and hardware configuration, so it is misleading to claim that data movement always dominates computation. The more useful observation is that modern processors have become so capable at arithmetic that, for certain workloads, improving the way data is accessed can produce a larger benefit than adding more conventional computing units. In-memory approaches target precisely this part of the architecture.

The terms processing-near-memory (PNM) and processing-in-memory (PIM) describe related approaches, but they should not be treated as interchangeable labels for one identical architecture. PNM generally places dedicated computational logic physically close to memory, reducing the distance that data must travel and potentially increasing available internal bandwidth. Three-dimensional packaging and high-bandwidth memory can provide useful building blocks for such designs, although HBM itself is a memory technology rather than a synonym for PNM. The processor and memory remain distinct components, but the architecture is designed to reduce the cost of communication between them.
PIM takes the concept further by placing some computational capability within the memory subsystem or integrating it closely with the memory array. Depending on the implementation, computation may be performed alongside memory banks, within specialized logic associated with those banks, or through operations that exploit the physical behavior of memory cells. Research literature also uses related terms such as compute-in-memory (CIM), and terminology is not completely standardized across hardware communities. The practical distinction is therefore better understood as a spectrum: some architectures move computation closer to memory, while others allow selected computations to occur directly within or alongside the structures that store the data.
One important implementation of in-memory computing uses memory arrays to perform highly parallel mathematical operations. In an analog crossbar, for example, programmable conductance values can represent numerical weights while input signals are applied across rows of the array. The resulting currents combine along columns according to the electrical properties of the circuit. Ohm's law and Kirchhoff's current law provide the physical basis for these operations, allowing many multiply-and-accumulate calculations to occur concurrently rather than requiring every operation to be explicitly executed by a separate digital arithmetic unit.
The architectural advantage is more specific than saying that the memory array "becomes a computer." A conventional accelerator may repeatedly fetch stored weights, move them into computational units, perform calculations, and return results through the memory hierarchy. A suitable in-memory design can keep some of those weights physically close to the computation, reducing repeated transfers between storage and arithmetic logic. The surrounding system still requires input movement, control, sensing, conversion, communication, and output processing. In other words, in-memory computing does not eliminate data movement altogether. Its purpose is to reduce particular forms of movement that would otherwise consume time, bandwidth, or energy.
Artificial intelligence has renewed interest in these architectures because neural networks often perform large numbers of repeated operations on structured collections of weights and activations. Matrix multiplication and related tensor operations can require substantial movement of model parameters and intermediate values, particularly when memory bandwidth becomes a limiting factor. This creates an opportunity for specialized hardware to keep frequently used data closer to the computational resources. In-memory approaches are therefore attractive not because AI automatically requires PIM, but because some AI workloads expose exactly the kind of data-access patterns that PIM is designed to address.
The potential benefit depends heavily on the model and implementation. Inference workloads with predictable dataflow and substantial weight reuse may be easier to map onto specialized memory-centric hardware than applications with irregular control flow or frequent synchronization. Training can introduce additional challenges because model parameters must be updated repeatedly and numerical behavior can be more demanding. Analog compute-in-memory approaches also have to account for signal precision and conversion overhead, while digital PIM designs face their own constraints involving memory organization and programming models. For these reasons, PIM should be viewed as a workload-specific optimization rather than a universal replacement for GPUs, CPUs, or conventional AI accelerators.

The strongest case for PIM generally appears when a workload spends substantial effort accessing data and relatively less time performing complex control-oriented computation. Large-scale graph processing is one example because graph algorithms can involve repeated accesses to distributed data structures with limited locality. Recommendation and database workloads can present similar challenges when the cost of retrieving and moving information becomes a significant part of execution. Certain AI inference tasks may also benefit when large collections of model parameters can remain close to the computation instead of repeatedly traveling through conventional memory interfaces.
There are equally important situations where PIM may provide little practical benefit. A workload that is already compute-bound may gain less from reducing memory traffic, while highly irregular algorithms can be difficult to map efficiently onto specialized memory structures. Applications requiring frequent communication between host processors and memory-side compute units can also lose some of the expected advantage through synchronization and coordination overhead. High numerical precision may further limit the usefulness of some analog approaches. These constraints make system-level evaluation essential. A PIM design should be judged by the performance, energy use, accuracy, programmability, and total system complexity of the complete workload rather than by the theoretical throughput of its memory array alone.
Reducing data movement does not come without additional hardware complexity. A PIM architecture may require dedicated logic, modified memory organizations, specialized interconnects, additional control mechanisms, or new methods for coordinating computation with conventional processors. Analog implementations introduce another set of considerations, including device variation, electrical noise, limited precision, signal conversion, and calibration. Digital implementations avoid some of these analog limitations but still have to determine which operations can be executed efficiently inside the memory subsystem and how those operations interact with the rest of the system.
Software is another major part of the equation. Mainstream operating systems, compilers, libraries, and machine-learning frameworks have been developed around conventional processor and accelerator models. A PIM system may require new programming abstractions, compiler support, workload partitioning strategies, and tools for deciding which operations should remain on the host processor and which should move toward memory. These requirements create a practical adoption barrier even when the underlying hardware demonstrates strong results. The most useful PIM architectures will therefore need more than efficient circuits; they will need programming environments that allow developers to exploit memory-side computation without redesigning an entire software stack for every application.
Performance claims for in-memory computing are particularly easy to misunderstand because different studies may measure very different parts of a system. An analog crossbar may demonstrate extremely efficient matrix operations, for example, while the complete accelerator still requires digital control, data conversion, memory management, and communication. Similarly, a PIM architecture may reduce traffic across one interface without eliminating transfers elsewhere in the system. Comparing a specialized memory array with a complete conventional accelerator can therefore produce a misleading impression unless both measurements use comparable workloads, precision, system boundaries, and performance targets.
A more useful evaluation considers the complete path from input to output. Questions include how much data must cross a conventional processor-memory interface, how much computation occurs locally, what conversion or synchronization overhead is introduced, and whether the resulting accuracy meets the application's requirements. It is also important to examine software development effort and the hardware required outside the memory subsystem. This broader perspective changes the question from "How fast is the PIM array?" to "How efficiently can the complete system perform a useful workload?" That distinction is essential when assessing research prototypes, commercial products, and future architectural proposals on equal terms.
In-memory computing and processing-in-memory do not represent a return to a computer architecture in which conventional processors disappear. Digital CPUs, GPUs, accelerators, caches, memory controllers, and storage systems remain essential because they provide flexibility that specialized memory-centric architectures cannot easily reproduce. A more realistic future is heterogeneous: different parts of a computing system perform the tasks for which they are best suited. Conventional processors can handle control-heavy operations, while memory-side logic can accelerate selected operations that would otherwise require extensive movement through the memory hierarchy.
The importance of PIM will ultimately depend on whether that specialization produces measurable system-level benefits. The technology is most compelling when a workload is strongly affected by memory access, when its computations can be expressed efficiently within the available memory architecture, and when the additional hardware and software requirements are justified by the result. That makes PIM less of a universal successor to conventional computing and more of an architectural tool for a specific class of problems. By reconsidering the relationship between storage and computation, these systems offer another way to address a fundamental limitation of modern hardware: sometimes the most expensive part of a calculation is not doing the work, but moving the information needed to do it.