High-Bandwidth Memory is not fast DRAM. It is DRAM reorganized into a vertical tower, wired through the silicon itself, and mounted millimeters from the processor. The cells are the same. The manufacturing process is not.
HBM achieves bandwidth through extreme parallelism and proximity, not clock speed. A single stack in HBM4 contains 256 banks per die, connected by thousands of vertical copper columns called through-silicon vias. The result is a device that delivers bandwidth by being wide and close, but costs three times the silicon and adds 19 materials engineering steps to an already complex DRAM process. HBM yields one-third the gigabytes per wafer compared to standard DRAM, which is why the memory industry is now the bottleneck in AI hardware production.
The stack: DRAM dies, a base die, and an interposer
An HBM device is a vertical assembly. Four to sixteen DRAM dies sit on top of a logic base die. The entire stack sits on a silicon interposer next to the GPU. The DRAM dies store bits using ordinary DRAM cells. The base die handles I/O, testing, and routing between the stack and the processor. The interposer is a passive silicon layer with wiring fine enough to route over 1,000 individual wires between the GPU and each HBM stack, a density no circuit board can achieve.
The DRAM dies are thinned to roughly 50 micrometers after fabrication, then stacked and bonded using microbumps with 30 to 50 micrometer pitch. Each die connects to the one below it through thousands of through-silicon vias. The base die connects to the interposer, which connects to the GPU. The entire assembly is encapsulated in molded underfill to manage mechanical stress and provide structural integrity.
Industry discussion now centers on increasing total cube thickness from 720 to 775 microns to accommodate more layers, according to Jaesik Lee, Vice President of Package Engineering at SK hynix America, speaking at Hot Chips 2026. The constraint is not just mechanical. Thinner dies mean less margin for TSV depth, tighter tolerances on warpage, and higher thermal density as more active silicon is compressed into the same footprint.
Through-silicon vias: the vertical wiring problem
Through-silicon vias are copper columns drilled straight through thinned silicon dies. They carry signals, power, and ground vertically through the stack. A single HBM device contains thousands of TSVs. They are what make vertical stacking electrically viable.
Manufacturing a TSV requires etching a trench into silicon, depositing an insulating liner, filling the trench with copper via electroplating, polishing the surface, thinning the wafer from the back side to reveal the via, and bonding the die to the next layer. Applied Materials reports that HBM adds approximately 19 materials engineering steps to the roughly 700 process steps required for conventional DRAM: 10 frontside steps for TSVs and interconnect pillars, 9 backside steps to reveal the vias and create backside interconnects.
The challenge is not just adding steps. TSVs consume silicon area that would otherwise store bits. HBM3 achieves 0.16 gigabits per square millimeter compared to 0.296 gigabits per square millimeter for SK Hynix D1z DDR4, an 85 percent density penalty. Sangwook Han, Design Team Member at Samsung Memory Business, explained the scaling problem at Hot Chips 2026:
Sangwook Han, Samsung Memory BusinessAll these dies are connected through tens of thousands of TSVs. While maximum capacity has increased successfully in each generation, the primary motivation remains bandwidth because it has essentially doubled with each generation. To achieve higher bandwidth, we face two major bottlenecks: TSVs between dies and the PHY I/O in the base die. TSV scaling is relatively straightforward, since increasing the speed of the TSVs is relatively difficult and sometimes risky. So we have been increasing the number of TSVs. The drawback of this approach is the silicon area taken by the TSV array, which forces us to continuously reduce the pitch of the TSVs.
Reducing TSV pitch means tighter process control on etch depth, liner uniformity, and copper fill. As the industry moves toward sub-5-micrometer TSV diameters for future generations, deposition uniformity variations become yield limiters. Applied Materials and imec established the industry's first TSV process flow and integration scheme approximately 15 years ago, forming the basis for current HBM TSV standards, but each generation requires new materials and tighter tolerances.
Parallelism: 256 banks and dual PHY architecture
HBM achieves bandwidth through parallelism, not speed. Where a typical DDR module might have 16 or 32 banks, HBM3E implements 128 banks per DRAM die and HBM4 doubles this to 256 banks per die. Each bank can operate independently. The result is hundreds of simultaneous data paths feeding into the base die.
Raghu Sreeramaneni, Fellow for HBM Design Architecture at Micron, described the architecture at Hot Chips 2026:
Raghu Sreeramaneni, MicronThe difference is in what kind of technology, packaging, and architecture is designed around that cell technology to be able to deliver higher bandwidth. The main thing is extremely high parallelism. There are way more banks that can operate in parallel. There are way more data paths that can take that data and connect it to the base die. HBM3E had 128 banks on every DRAM die. HBM4 goes all the way to 256 banks per DRAM die, and then, obviously, all the DRAM dies connect to the base die, which forms the interface to the GPU or XPU. There are essentially two PHYs. The base die has a PHY that talks to the XPU through the microbumps and interposer, and there's a TSV PHY that is connecting all the DRAMs to the base die.
The dual PHY design means the base die is doing complex, high-speed work at the bottom of the stack, generating heat in the location farthest from the cooling solution. The TSV PHY manages vertical communication through the stack. The external PHY drives the interface to the GPU across the interposer. Both operate at multi-gigabit speeds, and both contribute to the thermal problem.
Thermal extraction: heat flows up, cooling is at the top
Heat in an HBM stack is generated primarily at the base die and must escape through the DRAM dies above it to reach the heat sink at the top. DRAM cells require periodic refresh operations to retain data, and refresh rates increase with temperature. If the stack runs too hot, refresh overhead rises and reliability degrades.
Sreeramaneni described the thermal constraint at Hot Chips 2026:
Raghu Sreeramaneni, MicronIf you think of an HBM cube, the heat is getting extracted at the top. The base die is typically the hottest part of the die, because it's doing some of the most complex work. It has the highest-speed interconnections, so you're generating a lot of heat at the bottom. Your heat sink and cooling is at the very top. There is thermal resistance through the cube, and DRAM doesn't like to be hot. So refreshes of DRAM start to become a problem for reliability if it gets too hot. Figuring out what the thermal solutions are is very, very critical, and it's getting to the point where we are now architecting solutions around thermals instead of the other way around.
Packaging solutions include molded underfill materials with higher thermal conductivity, thermal TSVs that carry no signal but provide a vertical heat path, and tighter control of die warpage to maintain contact pressure across microbumps. Jaesik Lee noted at Hot Chips 2026 that adding layers increases power density and creates thermal challenges, especially as each die adds oxide layers that act as thermal insulators. Thermal management now drives architectural decisions.
Why HBM is a manufacturing bottleneck
HBM uses the same DRAM process technology as DDR, but the additional TSV and packaging steps reduce wafer output. Jim Handy, General Director at Objective Analysis, explained the supply constraint at Hot Chips 2026:
Jim Handy, Objective AnalysisIf you look at the number of gigabytes you get per wafer, it's a third as much as what you get with DDR, standard DRAM. What that means is, all of a sudden, HBM just is demanding a colossal number of wafers, and that's causing all DRAM to go into shortage because they're all made on the same product process lines.
The silicon footprint is also larger. An 8-stack HBM4 configuration with 12 layers per stack contains 12,344 square millimeters of memory-related silicon, more than eight times the area of a typical GPU die. The combination of lower wafer efficiency, higher silicon consumption per device, and shared fab capacity with standard DRAM creates a structural supply constraint.
TSV etching and electroplating tools are a specific bottleneck because they are incremental to the standard DRAM process and have long lead times. Back-end packaging capacity for thermo-compression bonding and molded underfill is another constraint. Yield losses from stacking defects compound the problem: a single bad TSV in one die can scrap the entire stack. HBM production scales more slowly than demand, and the memory industry has become the limiting factor in AI accelerator production.
Sources & further reading
- SemiEngineering, Issues Stack Up With More HBM Layers Read →
- SemiAnalysis, Scaling the Memory Wall: The Rise and Roadmap of HBM Read →
- Applied Materials, HBM: Materials Innovation Propels High-Bandwidth Memory Read →
- Siemens EDA, HBM3E & HBM4 IC Design Guide Read →
- 404k Research, Micron Technology on the AI Memory Bottleneck Read →