r/LocalLLaMA · · 1 min read

Micron's memory wall chart. Compute up ~3x every two years, HBM bandwidth under 2x

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Micron's memory wall chart. Compute up ~3x every two years, HBM bandwidth under 2x

From Raghu Sreeramaneni's memory tutorial at hot chips 2026. top line is normalised tflops for tpu v3 through r200, bottom line is hbm2e through hbm4, both log scale, so the distance between them is a lot wider than it looks.

The three boxes down the right are the fixes people are actually building. Memory beside the compute, memory closer on a shorter link, then multiply units inside the memory itself. Samsung has that last one shipping in lpddr5x and measured 3.01x tokens a second on llama 3.1 8B.

Full analysis (this slide sits in the memory chapter): https://allaboutchips.com/#memory

submitted by /u/Summit-Star001
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA