Tech & AI News
Hacker News

Cache-to-Cache: Direct Semantic Communication Between Large Language Models

Cache-to-Cache (C2C) enables direct semantic communication between Large Language Models by projecting and fusing KV-caches via a neural network and learnable gating mechanism. This approach bypasses intermediate text generation, achieving 6.4-14.2% higher accuracy than individual models. C2C also outperforms text-based communication by up to 5.4% while delivering a 2.5x speedup in latency.