Hacker News
A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation
FIBER introduces a “fiber” execution unit that decouples thread control from private registers, allowing dynamic parallelism and fine-grained register-level scheduling while sharing SM registers. Implemented via ISA, microarchitectural, and compiler extensions, it delivers 2.25× end-to-end speedup on Ampere and up to 2.49× kernel gains in mixed-precision LLM serving, with 1.8× and 2.09× improvements on Hopper and Blackwell respectively.