Tech & AI News
Hacker News

Why your local LLM feels dumber than it is

The author compares reproducible LLM inference across Nvidia H200, AMD B200, and RTX 6000 GPUs using three attention backends—FlashAttention 2, FlashInference, and Triton—on over 9,000 token positions from long prompts. Results show hardware-specific, byte-identical outputs that differ predictably due to bfloat16 rounding and arithmetic variations, revealing distinct relative divergences without a single “right” answer.