- Samsung's LPDDR5X-PIM achieved 3.01× higher token throughput than conventional LPDDR5X in Llama 3.1 8B inference tests at Hot Chips 2026.
- Digital PIM delivers 2.67 to 4.79× throughput gains over baseline analog PIM for 8-bit integer multiply operations, per Seoul National University research.
- Analog memristive crossbars reached 95.24% accuracy in Khalifa University's February 2025 Nature Electronics demonstration.
- Internal bandwidth hit 614 GB/s in Samsung's LPDDR5X-PIM, enabling parallel matrix operations across distributed processing units.
Framing
Why compute-in-memory splits into analog and digital camps
Analog compute-in-memory using memristive crossbar arrays demonstrated 95.24% accuracy for language processing tasks with 90% hardware reduction compared to conventional architectures, as reported in Nature Electronics on February 11, 2025, validating the approach for edge AI inference where power budgets justify trading bit-level precision for energy efficiency.
Compute-in-memory architectures perform matrix operations directly in SRAM, DRAM, or resistive arrays, eliminating the energy cost of shuttling weights and activations between separate memory and logic tiles. Analog approaches exploit physical laws for one-cycle dot products at the cost of precision limits from device mismatch and ADC quantization, while digital CIM embeds multiply-accumulate units inside memory arrays to preserve full precision at higher area overhead.
Processing-in-memory (PIM) paradigm is a promising solution to minimize the cost of data movement in the von Neumann architecture. DRAM-based PIM technology can be implemented in two ways: analog and digital. Although computing within memory cells without digital compute units at peripherals is certainly appealing, DRAM-based analog PIM architectures still have various limitations to overcome. In particular, we identify three major challenges to bring analog PIM closer to production: compatibility with commodity DRAM microarchitecture, reliability, and coverage of operations.
Seoul National University Department of Computer Science & Engineering, Ph.D. Dissertation, November 2023
The choice between analog and digital CIM depends on target operations, operand bit-width, and device specifications rather than one approach being universally superior. Digital processing-in-memory architectures achieve 2.67–4.79× higher throughput than analog PIM for 8-bit integer multiplication in commodity DRAM devices, according to Seoul National University research published in 2024, while analog PIM exhibits 4.96× to 8.48× higher throughput for bitwise operations in DDR4. For arithmetic operations in HBM2, digital PIM shows higher throughput across all bit-width configurations, illustrating the workload-dependent performance trade-offs. These precision and throughput differences directly affect quantization strategies for deploying neural networks on CIM hardware.
Samsung LPDDR5X-PIM demonstration
Samsung demonstrated LPDDR5X-PIM at Hot Chips 2026, achieving 3.01× higher token throughput than conventional LPDDR5X on Llama 3.1 8B inference with 16 GB per package and 614 GB/s internal bandwidth, though commercial availability and software ecosystem support remain unconfirmed.
Samsung's LPDDR5X-PIM architecture places compute logic directly inside each DRAM bank, eliminating the need to move data across the external memory bus for inference operations. Each chip contains 16 banks, with one processing-in-memory block per bank capable of executing matrix operations at 2.4 TOPS (SINT4) or approximately 1.2 TFLOPS (FP8) across 15 precision combinations from INT4 to FP8. The system delivers 614 GB/s of internal bandwidth between the PIM blocks and memory arrays, compared to 76.8 GB/s for conventional LPDDR5X external bandwidth. On Llama 3.1 8B inference, Samsung demonstrated 81.3 tokens per second versus 27 tokens per second for conventional LPDDR5X, a 3.01× improvement. Each package holds 16 GB. The architecture offloads matrix multiplications to the PIM blocks while the host processor handles control flow, keeping weight data local to the compute and transmitting only activations and results across the external interface.
Samsung LPDDR5X-PIM token throughput versus conventional LPDDR5X
Industry solutions in flight.
The approaches being pursued in parallel, with the status of each.
Samsung LPDDR5X-PIM
Samsung demonstrated LPDDR5X-PIM silicon at Hot Chips 2026, claiming 3.01× higher token throughput on Llama 3.1 8B inference with 614 GB/s internal bandwidth. Productization targeted for 2026; no volume shipments or customer deployments confirmed.
How the story evolved.
A dated log of material developments, each entry linked to its primary source.
Khalifa University demonstrated video processing on a self-calibrating analogue memristor array with on-chip calibration to compensate for device variability, enabling real-time edge inference. This advance addresses a key challenge in analog compute-in-memory systems: maintaining accuracy despite inherent analog device drift and mismatch.
Key players.
Mythic AI
Ships analog SRAM crossbar processors for edge AI inference in embedded vision and industrial IoT, encoding weights as conductances for one-cycle multiply-accumulate operations.
Outlook · what to watch.
The catalysts that will show whether this is closing, holding, or worsening.
Nature Electronics memristor array video processing
Nature Electronics publishes video processing results on self-calibrating analogue memristor arrays. Demonstrates analog compute-in-memory capability for real-time inference workloads beyond matrix operations.
IBM PCM-based edge inference at IEDM
IBM Research presents phase-change memory compute-in-memory results for edge inference at IEEE IEDM. Extends analog in-memory computing from lab to practical edge deployment scenarios.
IBM analog MoE computing in Nature Computational Science
IBM Research publishes cover feature in Nature Computational Science on analog in-memory computing for mixture-of-experts models. Validates analog approach for large-scale AI inference, challenging digital-only assumptions.
Samsung LPDDR5X-PIM demonstration
Samsung demonstrates LPDDR5X-PIM at Hot Chips 2026. Brings processing-in-memory to mainstream DRAM, signaling industry shift toward memory-centric compute for inference acceleration.
Ankhdjet ternary compute-in-ROM compiler
Ankhdjet open-source compiler for mask-programmed ternary compute-in-ROM released on arXiv. Enables low-cost inference on fixed-function memory arrays, expanding accessibility of compute-in-memory techniques.