bitdepth.co
LIVE
Pillars/Silicon · Architecture/In-memory computing: analog and digital compute-in-memory for AI inference
Active pillar · Updated weekly · Edition 1

In-memory computing: analog and digital compute-in-memory for AI inference

Tracks compute-in-memory architectures that perform matrix operations directly in memory arrays, eliminating data movement overhead.

Edited byBitDepth Team
First publishedSep 8, 2026
Last revisedSep 8, 2026 · 19:45 UTC
  • Samsung's LPDDR5X-PIM achieved 3.01× higher token throughput than conventional LPDDR5X in Llama 3.1 8B inference tests at Hot Chips 2026.
  • Digital PIM delivers 2.67 to 4.79× throughput gains over baseline analog PIM for 8-bit integer multiply operations, per Seoul National University research.
  • Analog memristive crossbars reached 95.24% accuracy in Khalifa University's February 2025 Nature Electronics demonstration.
  • Internal bandwidth hit 614 GB/s in Samsung's LPDDR5X-PIM, enabling parallel matrix operations across distributed processing units.
1

Framing

Why compute-in-memory splits into analog and digital camps

Analog compute-in-memory using memristive crossbar arrays demonstrated 95.24% accuracy for language processing tasks with 90% hardware reduction compared to conventional architectures, as reported in Nature Electronics on February 11, 2025, validating the approach for edge AI inference where power budgets justify trading bit-level precision for energy efficiency.

Compute-in-memory architectures perform matrix operations directly in SRAM, DRAM, or resistive arrays, eliminating the energy cost of shuttling weights and activations between separate memory and logic tiles. Analog approaches exploit physical laws for one-cycle dot products at the cost of precision limits from device mismatch and ADC quantization, while digital CIM embeds multiply-accumulate units inside memory arrays to preserve full precision at higher area overhead.

Processing-in-memory (PIM) paradigm is a promising solution to minimize the cost of data movement in the von Neumann architecture. DRAM-based PIM technology can be implemented in two ways: analog and digital. Although computing within memory cells without digital compute units at peripherals is certainly appealing, DRAM-based analog PIM architectures still have various limitations to overcome. In particular, we identify three major challenges to bring analog PIM closer to production: compatibility with commodity DRAM microarchitecture, reliability, and coverage of operations.

Seoul National University Department of Computer Science & Engineering, Ph.D. Dissertation, November 2023

The choice between analog and digital CIM depends on target operations, operand bit-width, and device specifications rather than one approach being universally superior. Digital processing-in-memory architectures achieve 2.67–4.79× higher throughput than analog PIM for 8-bit integer multiplication in commodity DRAM devices, according to Seoul National University research published in 2024, while analog PIM exhibits 4.96× to 8.48× higher throughput for bitwise operations in DDR4. For arithmetic operations in HBM2, digital PIM shows higher throughput across all bit-width configurations, illustrating the workload-dependent performance trade-offs. These precision and throughput differences directly affect quantization strategies for deploying neural networks on CIM hardware.

3

Samsung LPDDR5X-PIM demonstration

Samsung demonstrated LPDDR5X-PIM at Hot Chips 2026, achieving 3.01× higher token throughput than conventional LPDDR5X on Llama 3.1 8B inference with 16 GB per package and 614 GB/s internal bandwidth, though commercial availability and software ecosystem support remain unconfirmed.

Samsung's LPDDR5X-PIM architecture places compute logic directly inside each DRAM bank, eliminating the need to move data across the external memory bus for inference operations. Each chip contains 16 banks, with one processing-in-memory block per bank capable of executing matrix operations at 2.4 TOPS (SINT4) or approximately 1.2 TFLOPS (FP8) across 15 precision combinations from INT4 to FP8. The system delivers 614 GB/s of internal bandwidth between the PIM blocks and memory arrays, compared to 76.8 GB/s for conventional LPDDR5X external bandwidth. On Llama 3.1 8B inference, Samsung demonstrated 81.3 tokens per second versus 27 tokens per second for conventional LPDDR5X, a 3.01× improvement. Each package holds 16 GB. The architecture offloads matrix multiplications to the PIM blocks while the host processor handles control flow, keeping weight data local to the compute and transmitting only activations and results across the external interface.

Samsung LPDDR5X-PIM token throughput versus conventional LPDDR5X

Tokens per second on Llama 3.1 8B inference and internal bandwidth (GB/s)
Supply
Demand
Gap
100755025202220232024202520262027PROJTODAY3.01× throughput improvement81.3 vs. 27 tokens/second
Read: Samsung LPDDR5X-PIM delivered 81.3 tokens per second on Llama 3.1 8B inference versus 27 tokens per second for conventional LPDDR5X, enabled by 614 GB/s internal PIM bandwidth compared to 76.8 GB/s external DRAM bandwidth. The implication: Processing-in-memory architectures can triple inference throughput by eliminating external memory bottlenecks, moving compute to where the model weights reside rather than shuttling gigabytes of data across a narrow bus.
§ 03 · Tracker

Industry solutions in flight.

The approaches being pursued in parallel, with the status of each.

SOL — 01

Samsung LPDDR5X-PIM

In development

Samsung demonstrated LPDDR5X-PIM silicon at Hot Chips 2026, claiming 3.01× higher token throughput on Llama 3.1 8B inference with 614 GB/s internal bandwidth. Productization targeted for 2026; no volume shipments or customer deployments confirmed.

§ 04 · Timeline

How the story evolved.

A dated log of material developments, each entry linked to its primary source.

February 2025RESEARCH
Khalifa University demonstrates real-time video processing on analog memristor array

Khalifa University demonstrated video processing on a self-calibrating analogue memristor array with on-chip calibration to compensate for device variability, enabling real-time edge inference. This advance addresses a key challenge in analog compute-in-memory systems: maintaining accuracy despite inherent analog device drift and mismatch.

Nature Electronics
§ 05 · Map

Key players.

Analog compute-in-memory · US

Mythic AI

Ships analog SRAM crossbar processors for edge AI inference in embedded vision and industrial IoT, encoding weights as conductances for one-cycle multiply-accumulate operations.

§ 06 · Forward look

Outlook · what to watch.

The catalysts that will show whether this is closing, holding, or worsening.

2025 · ongoing

Nature Electronics memristor array video processing

Nature Electronics publishes video processing results on self-calibrating analogue memristor arrays. Demonstrates analog compute-in-memory capability for real-time inference workloads beyond matrix operations.

Watch: Publication date confirmation and performance metrics versus baseline GPU inference.
· Q4

IBM PCM-based edge inference at IEDM

IBM Research presents phase-change memory compute-in-memory results for edge inference at IEEE IEDM. Extends analog in-memory computing from lab to practical edge deployment scenarios.

Watch: IEDM proceedings availability and power efficiency figures for edge models.
· ongoing

IBM analog MoE computing in Nature Computational Science

IBM Research publishes cover feature in Nature Computational Science on analog in-memory computing for mixture-of-experts models. Validates analog approach for large-scale AI inference, challenging digital-only assumptions.

Watch: Citation adoption by other labs and replication of MoE efficiency claims.
2026 · Q3

Samsung LPDDR5X-PIM demonstration

Samsung demonstrates LPDDR5X-PIM at Hot Chips 2026. Brings processing-in-memory to mainstream DRAM, signaling industry shift toward memory-centric compute for inference acceleration.

Watch: Availability timeline, power-per-inference metrics, and OEM adoption announcements.
2026 · Q3

Ankhdjet ternary compute-in-ROM compiler

Ankhdjet open-source compiler for mask-programmed ternary compute-in-ROM released on arXiv. Enables low-cost inference on fixed-function memory arrays, expanding accessibility of compute-in-memory techniques.

Watch: Peer review acceptance and adoption by inference framework projects.

Industry standard for compute-in-memory interfaces

We use cookies

We use cookies to ensure you get the best experience on our website. For more information on how we use cookies, please see our cookie policy.