VentureBeat
Nvidia’s new technique cuts LLM reasoning costs by 8x without losing accuracy

Nvidia's new technique, Dynamic Memory Sparsification (DMS), reduces large language model reasoning costs by up to 8x. DMS compresses the key value cache, allowing models to "think" longer without degrading accuracy. This technique enables more efficient memory use.