I tested KV cache prefetch on 744,000 real agent tool calls. The prediction does not pay.
Open RFCs in vLLM and SGLang propose predicting when an agent returns from a tool call and prefetching its KV cache. Tested on 744,000 real tool calls and a real vLLM server, the prediction does not pay. Host tier size, the store path and one scheduler flag move the tail.