Business Tech

Micron HBM3E: features, specifications and use cases

Micron HBM3E is high-bandwidth memory designed for AI accelerators, high-performance computing systems and other processors that need far more memory throughput than conventional server DRAM can provide. Micron offers the technology in 24GB 8-high and 36GB 12-high stacks, with both configurations delivering more than 1.2TB/s of bandwidth per stack.

HBM3E achieves that bandwidth by stacking DRAM dies vertically and connecting them through through-silicon vias, or TSVs. The memory then sits very close to the GPU or accelerator inside an advanced package rather than occupying conventional DIMM sockets elsewhere on the server motherboard.

As of August 2026, HBM3E is no longer Micron’s newest HBM generation. The company started high-volume production of HBM4 in early 2026. However, HBM3E remains an important current technology powering AI hardware including NVIDIA H200 and AMD Instinct MI350-series accelerators.

Micron HBM3E comes in 24GB and 36GB stacks

Micron’s first HBM3E configuration combines eight DRAM dies into a 24GB 8-high stack. Its higher-capacity version increases the stack to 12 dies and 36GB.

Both provide a pin speed above 9.2Gbps and more than 1.2TB/s of bandwidth per placement. The interface uses 1,024 I/O pins, 16 channels and 32 pseudo channels.

SpecificationMicron HBM3E
Memory typeHigh Bandwidth Memory 3E
Capacities24GB, 36GB
Stack configurations8-high, 12-high
DRAM processMicron 1β (1-beta)
Pin speedMore than 9.2Gbps
BandwidthMore than 1.2TB/s per stack
I/O interface1,024 pins
Channels16
Pseudo channels32
Burst length8
Package dimensions11 × 11 × 0.72mm*
Operating temperature0°C to 105°C
Packaging supportCoWoS and SiP
Main applicationsAI training, AI inference and HPC

*Micron’s published product brief lists these dimensions for its HBM3E product. Specific system implementations depend on the accelerator and packaging design.

How high-bandwidth memory works

Conventional server DRAM generally connects to a processor through memory channels leading to DIMMs installed on the motherboard. Increasing memory bandwidth through that architecture requires faster signalling, additional memory channels or both.

HBM takes a different approach. Multiple DRAM dies are stacked vertically and connected through TSVs, while an extremely wide memory interface allows a large amount of data to move simultaneously.

Micron says its HBM3E design uses twice as many TSVs as the HBM3 products referenced in its product documentation. Its advanced packaging supports technologies including chip-on-wafer-on-substrate and system-in-package designs.

The result is very high aggregate bandwidth while keeping the memory physically close to the processor consuming the data.

That proximity is particularly valuable for AI accelerators. Thousands of GPU or accelerator cores cannot remain productive if they spend too much time waiting for model weights, activations and other data to arrive from memory.

Why does AI need HBM3E?

Modern AI accelerators can perform enormous numbers of calculations every second. Supplying those compute units with enough data has therefore become one of the major system-design challenges.

HBM3E addresses two parts of that problem: bandwidth and capacity.

The bandwidth determines how quickly data can move between memory and the accelerator. Capacity determines how much of a model and its working data can remain close to the accelerator without being transferred to another GPU or system memory.

Micron positions HBM3E for both AI training and inference. Higher capacity can allow larger models or larger portions of a model to remain on an accelerator, while greater bandwidth helps feed the compute engines performing the actual matrix and tensor operations.

This does not mean HBM3E by itself makes an AI system fast. GPU architecture, interconnects, software, model design, precision, storage and the number of accelerators all affect overall performance.

Micron uses its 1-beta DRAM process

Micron manufactures HBM3E using its 1β, or 1-beta, DRAM process technology.

The company credits this process, together with CMOS, thermal and packaging improvements, for HBM3E’s combination of bandwidth and power efficiency. Micron also states that its 24GB 8-high HBM3E provides 50% greater memory capacity within the same stack height compared with the previous capacity point used in its comparison.

Power consumption is particularly important in AI data centres because HBM sits alongside accelerators that already consume substantial amounts of electricity.

Micron claims its 8-high and 12-high HBM3E products can consume up to 30% less power than competing HBM3E offerings. This is a manufacturer comparison rather than an independently established figure covering every competing configuration.

Micron also claims approximately 2.5 times the performance per watt of its previous-generation HBM2E. Its documentation says that figure is based on its design and simulation measurements.

Reliability features target large AI systems

Memory errors become increasingly significant as AI systems scale to hundreds or thousands of accelerators.

Micron therefore includes several reliability, availability and serviceability features in HBM3E. Its product documentation lists Reed-Solomon on-die error correction, soft and hard memory-cell repair, and automatic error check and scrub functionality.

The memory also includes programmable built-in self-test functionality that can generate system-representative traffic at full specification speed. Micron positions this as a way to improve testing and debugging during accelerator development and system qualification.

These functions matter less to an end user than to companies designing accelerators and servers, but they help explain why HBM is substantially more complex than simply stacking ordinary DRAM chips.

NVIDIA H200 uses Micron 24GB HBM3E

Micron began volume production of its 24GB 8-high HBM3E in 2024, with the memory qualified for NVIDIA’s H200 Tensor Core GPU.

NVIDIA specifies the H200 with 141GB of HBM3E capacity and 4.8TB/s of total GPU memory bandwidth. That was a major increase over the HBM3 memory configuration used by the earlier H100.

Micron’s own H100-versus-H200 testing platform used six 24GB HBM3E placements, giving 144GB of physical memory capacity in the tested configuration and 4.8TB/s of system memory bandwidth.

The distinction between per-stack and per-GPU bandwidth is important. Micron’s specification of more than 1.2TB/s describes an individual HBM3E stack. The total bandwidth available to a GPU depends on how many stacks it uses and how the accelerator’s memory interface is configured.

36GB 12-high HBM3E powers AMD Instinct MI350

The 12-high version increases capacity to 36GB per stack without changing the fundamental HBM3E interface.

AMD selected Micron’s 36GB HBM3E for its Instinct MI350-series accelerator platform. An MI350-series GPU provides 288GB of HBM3E and up to 8TB/s of total memory bandwidth.

Eight 36GB placements account for that 288GB capacity.

The larger capacity is useful for AI because fitting more model data into accelerator-local memory can reduce transfers between GPUs or between an accelerator and host memory. Such transfers can add latency and consume interconnect bandwidth.

Micron also points to scientific simulation, computational modelling and other HPC workloads as targets for the same memory technology.

HBM3E use cases

AI training

Training large language and multimodal models involves repeatedly moving enormous volumes of model parameters and intermediate data between memory and accelerator cores.

More than 1.2TB/s from each HBM3E stack helps increase the amount of data available to those cores. Higher capacity can also let accelerators work on larger datasets or model sections without as much communication with other devices.

Micron reports modelling that showed training-time improvements of more than 30% in its HBM3E-related comparisons. Those results are based on Micron modelling and referenced H100-class systems, so they should not be interpreted as a universal performance increase for every AI model.

AI inference

Inference presents a related memory challenge. Large models must move their weights through the accelerator while generating responses, and memory bandwidth can become a bottleneck even when the GPU itself has sufficient compute performance.

More memory capacity can also allow larger models or context workloads to stay resident on the GPU rather than relying on slower host-memory transfers.

NVIDIA specifically positions the H200’s larger and faster HBM3E memory as an advantage for generative AI and large-language-model inference.

High-performance computing

HBM3E is not limited to generative AI.

Scientific modelling, engineering simulation, weather and climate research, computational science and other HPC workloads can process very large datasets. These applications can also become limited by how quickly data reaches the processor rather than by raw arithmetic capability alone.

Micron therefore targets HBM3E at both AI accelerators and supercomputing systems.

HBM3E versus conventional DDR5 memory

HBM3E and DDR5 solve different problems.

DDR5 DIMMs provide relatively economical, expandable system memory for CPUs. Servers can install large capacities through multiple memory sockets, and individual modules can be replaced or upgraded.

HBM3E sits in the accelerator package and provides dramatically greater memory bandwidth. However, it is a far more specialised technology and is not an upgradeable module that a server administrator inserts into a DIMM slot.

Micron has previously described an individual HBM3E stack as providing more than 20 times the bandwidth of a standard DDR5-based server DIMM in its comparison.

That does not make DDR5 obsolete. AI servers commonly need both: HBM next to the GPU for highly bandwidth-sensitive accelerator workloads and large pools of DDR memory connected to the CPUs.

HBM4 is already succeeding HBM3E

HBM3E remains commercially important in 2026, but Micron has already moved its leading HBM technology to the next generation.

The company began volume shipments of 36GB 12-high HBM4 during the first quarter of 2026 for NVIDIA’s Vera Rubin platform. Micron HBM4 exceeds 11Gbps per pin and 2.8TB/s per stack.

That is approximately 2.3 times the bandwidth of Micron’s equivalent 36GB 12-high HBM3E, while Micron also claims more than a 20% improvement in power efficiency. HBM4 doubles the interface width to 2,048 I/O connections.

The move does not make HBM3E irrelevant. HBM generations remain tied to specific accelerator architectures, and existing H200 and MI350-based servers will continue using HBM3E throughout their operating lives.

HBM3E should therefore be viewed as a current production technology serving major AI platforms, while HBM4 takes over the highest-performance tier of Micron’s portfolio.

What Micron HBM3E means for businesses

HBM3E is not something a normal business purchases separately to upgrade a server. It arrives as part of an accelerator such as an NVIDIA H200 or AMD Instinct MI350-series GPU.

For an organisation evaluating AI infrastructure, the important specifications are therefore the total accelerator memory capacity and aggregate bandwidth, rather than simply whether the memory carries an HBM3E label.

Higher HBM capacity can help when deploying large models, while higher bandwidth can improve accelerator utilisation for memory-intensive workloads. However, the entire platform still matters, including GPU compute capability, interconnect architecture, host memory, networking, cooling and software support.

For South African organisations, that makes HBM3E most relevant through imported AI servers, cloud GPU services, research computing and enterprise AI infrastructure rather than through conventional component retail.

Micron HBM3E’s significance is ultimately straightforward: modern accelerators need enormous amounts of data delivered at extremely high speed, and vertically stacked memory provides the bandwidth needed to keep those processors working.