VentureBeat
Nvidia finds that simple linear math can replace costly AI model handoffs

Nvidia researchers introduced a cross-model KV-cache transfer method that linearly maps the prefilled cache from one LLM to another, eliminating the need to recompute the entire prefill stage when switching models. Experiments on compatible model families show the technique runs 2.7–25× faster while preserving up to 98% of the target model’s standalone accuracy, cutting compute cost and latency in long-horizon, multi-LLM workflows.