Tech & AI News
Hacker News

What happens when a GPU reads memory

The article traces a single global-load instruction (LDG.E) in a vector-add kernel on an RTX 4090, detailing how the warp’s 32 lanes fetch 4-byte elements from global memory through the register file, operand collector, load/store unit, L1 cache, L2 slices, address translation, crossbar, and DRAM when misses occur. Timing experiments reveal the address read adds one cycle, while a shared-memory load from a register takes 24 cycles, illustrating the full hardware journey of the load.