Inference · 22 Aug 2026
Nvidia moves a KV cache between models with linear regression, skipping the re-prefill
Transferring a 32,768-token cache from Qwen3 14B to 32B took 278 milliseconds against seven seconds to recompute it, holding 73% to 98% of standalone accuracy.