I tested KV cache prefetch on 744,000 real agent tool calls. The prediction does not pay.

When a coding agent calls a tool, its KV cache sits in GPU memory until the agent comes back. Several research groups and open RFCs in vLLM and SGLang propose to predict the return and move the cache out and back in around it. I tested that on 744,000 real tool calls and on a real vLLM server. The prediction does not pay. What moves the tail is the size of the host memory tier, the store path, and one scheduler flag.

One tool call, three fates of the KV cacheThe agent sends a prompt, calls a tool, and returns. If its KV cache stayed on the GPU it waits 40 ms for the first token. If the cache was moved to host memory it waits 210 ms to reload. If the cache was dropped it waits 3.66 s to recompute a 32K-token context.prompttool runs, agent waitsagent returnsGPU memory24 GB, the scarce oneHost memory (DRAM)cheap, 0.21 s away for 32K tokensthis agent'sKV cachewait for the first token after the return0 ms
Figure 1. Three fates of a returning agent's cache, to scale on the red bar. A host-memory reload costs about one seventeenth of a recompute. That ratio is the reason the prediction idea fails: once a host tier exists, the most a clever policy can save is the difference between the first two scenarios.

The question

A coding agent spends much of its wall-clock time waiting on tools: test runs, builds, package installs. While it waits, its KV cache occupies the scarcest resource on an inference server, GPU memory. The open proposals are built on a simple idea: when a tool call starts, move the cache to host memory; shortly before the agent returns, bring it back; time the return by a prediction from past tool durations.

I had a supply-chain version of the same idea. Tool durations might cluster around a typical length, the way deliveries do, so the probability of return would rise toward that length and a serving system could learn, per tool type, when to demote and when to prefetch. Two questions followed. Can the prediction be good enough to pay off? And if it cannot, what actually decides how long an agent waits for its next token after a tool call?

What I did

  1. Measured the pause distributions on 743,819 real tool calls from Claude Code and Codex. The data is TraceLab, released by the University of Washington under CC BY 4.0: 8,058 sessions from 71 developers, with tool names and timings but no arguments.
  2. Measured the cost of every operation on one RTX 4090 running vLLM 0.30 with Llama 3.1 8B in FP8: prefill, decode, store to host memory, reload.
  3. Compared retention and prefetch policies in a discrete-event simulator built from those two inputs: 636 runs over load, host tier size, context length and hardware constants.
  4. Checked the simulator against the real engine, fourteen one-hour runs, and used those runs to test what moves the tail.

Every number below comes from a script or a table in the project.

1. Tool pauses have a falling hazard

Most pauses are short and the long ones hold the time. 70 percent of Claude Code pauses end within a second; the 4.9 percent of tool calls over a minute hold 92 percent of all tool time. Human approval waits are 37 percent of Claude's tool time on their own.

The shape of the long pauses decides whether prefetch can be timed. The idea I started with needs a return rate that rises toward a typical duration. The data shows the opposite. Pick how long the agent has already waited:

Expected remaining wait: the idea versus the dataThe idea I started withtool runs have a typical length, so the wait shrinksillustration, no dataexpected time still to waitshrinksWhat 574,485 pauses showClaude Code test runs, mean residual lifemeasuredexpected time still to wait54 safter 4.4 s of waiting
Figure 2. The longer a pause has lasted, the longer it is expected to last: 54 s left after 4 s, 113 s after 25 s, 361 s after 146 s. Every tool class behaves this way. The hazard falls with a log-log slope between −0.56 and −0.74 across 18 log-spaced bins, with no significant rise anywhere. There is no moment at which a return becomes likely.

The only sharp structure in the data is harness deadlines. Codex polls a running command at its yield deadline of 1, 5, 10 or 30 seconds, so a three-minute test reaches the engine as a chain of short pauses separated by small LLM steps; 30 percent of Codex pauses end on such a deadline. Claude Code pauses spike at its 120 and 600 second Bash timeouts and at sleeps the agent writes itself. A gateway that parses the tool call sees these values directly.

Better keys do improve the forecast. Conditioning the remaining-time distribution on tool class, command skeleton and developer, with shrinkage toward the parent level when a key has few observations, cuts the upper-quantile error for Claude by 27 to 37 percent, and the shrinkage works where Continuum's rule of a hundred observations per tool stays at the global baseline. At the low quantiles that a prefetch rule uses when stalls are expensive, the key barely matters. A better forecast never turns into a serving gain, for the reason in the next section.

2. Prediction-timed retention and prefetch do not beat LRU plus a host tier

The simulator replays TraceLab sessions closed-loop against a model of continuous batching, an HBM token pool, an LRU host tier with write-through stores and one shared PCIe link, calibrated on twelve single-session measurements. Contexts are scaled so the 99th-percentile session peaks at 32K tokens; load is counted in multiples of the eight sessions whose KV just fits the GPU pool. Six policies ran on it:

  • P0, no host tier at all.
  • P1, LRU with a host tier, which is what vLLM plus LMCache does out of the box.
  • P2, Continuum's time-to-live rule, implemented from the paper's equation with per-tool duration distributions.
  • P4-D, a deadline-aware prefetch that reloads just before a known harness deadline. Deadlines are exact here, which makes it an upper bound.
  • P5, an oracle that knows every return time.
  • P5-B, a strict Belady oracle: it never pins, evicts the session that returns latest, and prefetches each session just before its return.
p99 time to first token after a tool call, relative to plain LRU with a host tiermore than 10% better than LRU0.60.81.0, same as LRU1.21.4×1 load, 8 sessionsLRU p99: 0.31 s×2 load, 16 sessionsLRU p99: 4.13 s×4 load, 32 sessionsLRU p99: 16.6 sP2 Continuum TTL, ×1: 0.94P4-D deadline prefetch, ×1: 0.77P5 oracle, ×1: 0.68P2, ×2: 1.03P4-D, ×2: 1.04P5, ×2: 0.99P2, ×4: 1.10P4-D, ×4: 0.84P5, ×4: 0.80
P2, Continuum TTLP4-D, prefetch at known deadlinesP5, oracle that knows every return
Figure 3. Each dot is a policy's p99 divided by LRU's at the same load, simulator, 4090 constants, 20 GB host tier, mean of five seeds. Only the lightest load, where the GPU is idle enough to hold everything, shows gains inside the band. At ×4 the seed ranges are 11.2 to 19.6 s for LRU and 10.4 to 15.6 s for the oracle: they overlap. Without a host tier (P0) p99 is 4.3, 8.2 and 27.5 s.

The pre-registered success rule was a gain of more than 10 percent on p99 over the TTL policy, on all traffic or on the deadline subset, or a 10 percent gain in sessions served within a one-second SLO. Over 32 stable configurations:

  • Continuum's TTL rule pins the cache for 3.4 percent of tool pauses on this trace and behaves like LRU. The trace's memoryfulness coefficient is negative (−0.35), a case the Continuum paper notes as possible but did not observe on its own workloads.
  • Deadline prefetch clears 10 percent in 6 configurations, never in every seed. On the deadline subset, 9.
  • The Belady oracle clears 10 percent in 9, 5 of them in every seed, most of them at the lightest load. Its median p99 ratio to the TTL policy is 0.96.
  • The goodput rule never passes.
  • Forcing the memoryfulness coefficient positive makes the TTL pin up to two thirds of tool pauses. The TTL policy stays within about 10 percent of LRU at every setting.

The reason is Figure 1. A pin can save at most one reload plus the queueing it avoids, and with a host tier present that rarely outweighs the GPU capacity the pin holds. The PCIe link runs at 1 to 3 percent utilisation; its queueing delay comes from large stores of cold prompts, not from bursts of reloads. The negative result covers KV placement on a PCIe-attached host tier with that tier present. Remote or cross-node tiers, setups with no host tier, and uses of prediction for scheduling or admission were not tested.

3. What moves the tail

Engine evidence first. Every number in this section comes from a one-hour replay of real TraceLab sessions against vLLM 0.30 with LMCache on the 4090, with prompts of the right lengths and session-consistent prefixes so caching and eviction behave as on real traffic.

Change in the tail of post-tool TTFT per lever, measured on vLLM−80%−60%−40%−20%0+20%Host tier size, 4 to 8 GB16 sessions, seed 016 sessions, seed 116 sessions, seed 232 sessionsStore path, native vs LMCache16 sessions (p95)32 sessionsPrefill chunk cap16 sessions, seed 016 sessions, seed 116 sessions, seed 2−71%: 288 to 84 ms−39%: 112 to 69 ms−29%: 108 to 77 ms−44%: 294 to 165 ms−41% on p95−1%: tail unchanged, median halved−17%−35%−24%−71−39−29−44−41−1−17−35−24
Figure 4. Change in the mean of the slowest 5 percent of post-tool requests (p95 where marked), lever on versus off, on the real engine. Left is better. The host tier is the only lever that helps on every run. Not shown: the prefill cap at 32 sessions, where the GPU pool is oversubscribed, makes p95 13 percent worse and the median 2.3 times worse.

The host tier, sized well above the GPU pool

This machine cannot pin more than 8 GB without swapping, so I scaled the workload down five times and let 4 and 8 GB stand for host tiers at 1.9 and 3.7 times the GPU KV pool, the ratios that 20 and 40 GB would have at full scale. Doubling the tier cut the tail mean of post-tool TTFT by 71, 39 and 29 percent on three seeds at 16 sessions, and by 44 percent at 32.

What the slowest requests are made of follows that ratio. Each square below is one of the 100 slowest requests in a run:

What the slowest 5 percent of requests are made of, by host tier sizerecomputed from scratch in every runrecomputed in some runs, served from the tier in othersreloaded from the tier or queued, never recomputed
Figure 5. Share of the slowest 5 percent that are full recomputes of their context, across the engine runs at each tier size. At 0.75× the pool one long session evicted between steps can own the p99: in one baseline run, 8 of the 13 slowest requests were steps of a single 60K-token session, each recomputed from scratch in about 10 seconds.

The store path

LMCache's in-process connector writes the KV through to host memory inside the engine's step, so each forward pass waits for the copy. vLLM's own offloading connector (--kv-offloading-backend native) copies on a separate CUDA stream.

Batch 1, median of 51K tokens8K32K
Cold prefill, no connector71 ms571 ms3,648 ms
Cold prefill with LMCache store77 ms (+7%)649 ms (+14%)4,019 ms (+10%)
Cold prefill with native store71 ms (0%)569 ms (0%)3,670 ms (+1%)
Reload from host memory, LMCache21 ms63 ms210 ms
Reload from host memory, native36 ms125 ms409 ms

Under load the blocking store adds up. At 32 sessions LMCache stored 2.7 TiB during the hour and held the GPU for 238 seconds of it, 6.6 percent. Native copied a similar volume overlapped with compute. On the engine, native cut p95 by 41 percent at 16 sessions; at 32 sessions it halved the median and left the tail where it was. Its reloads run at about half LMCache's speed and did not show up in the tail at either load. On vLLM 0.30 the native connector is the safer default. Why the tail at 32 sessions does not improve is an open question: reload speed and store cost are both ruled out by the simulator.

The prefill chunk cap

long_prefill_token_threshold caps how many tokens one request may prefill per engine step, so short prefills queued behind a long one still get budget. At 16 sessions it improved the tail mean on all three seeds, by 17 to 35 percent, with p95 mixed. At 32 sessions the GPU pool is oversubscribed, 37 percent of post-tool requests recompute their context, and the cap backfires: p95 up 13 percent, the median 2.3 times worse. With the cap on, the engine keeps more prefills in flight at once, GPU KV usage rises from 46 to 53 percent, and fewer requests are served from the GPU cache. Set it last, and only with memory headroom.

What an operator should do, in order

OrderLeverDecisionEngine evidence
1Host tier sizeSize it well above the GPU KV pool: 3.7× rather than 1.9×.Tail mean −71, −39, −29% on three seeds at 16 sessions; −44% at 32.
2Store pathOn vLLM 0.30, use the native offloading connector rather than LMCache's in-process connector.p95 −41% at 16 sessions; median halved at 32; store penalty +1% against +10% on a cold 32K prefill.
3Prefill chunk capSet long_prefill_token_threshold only when the GPU pool has headroom.Tail mean −17 to −35% at 16 sessions; p95 +13% and median ×2.3 when oversubscribed.
4Retention timing, prediction-timed prefetchLeave to LRU once a host tier exists.Nothing on the engine. In simulation, within seed noise.

4. What the simulator can and cannot predict

Before any lever result came in, I fixed a pass rule: for a lever at a given load and seed, the engine's and the simulator's ratios (lever on over off) must agree within 0.15 on p95 and on the tail mean, absolute values within 15 percent or 30 ms, and direction must match when the predicted effect exceeds 15 percent. Three of ten checks passed.

Lever effect on the tail mean, engine against simulator, same seed00.51.0 (no effect)1.5Cap, 16 sessions, seed 0Cap, 16 sessions, seed 1Cap, 16 sessions, seed 2Cap, 32 sessionsHost tier, 16 sessions, seed 0Host tier, 16 sessions, seed 1Host tier, 16 sessions, seed 2Host tier, 32 sessionsNative connector, 16 sessionsNative connector, 32 sessionssimulator 0.76simulator 0.92simulator 0.36simulator 0.93simulator 0.19simulator 0.43simulator 0.57simulator 0.44simulator 1.38simulator 0.77engine 0.83, passengine 0.65, failengine 0.76, failengine 1.13, failengine 0.29, passengine 0.61, fail by 0.18engine 0.71, passengine 0.56, ratio within band, failed the absolute check by 5 msengine 0.70, failengine 0.99, fail
enginesimulator, same seedengine ± 0.15, the pass band
Figure 6. Tail-mean ratio per check. The host tier, the one large effect, lands in the band on three of four runs and on the right side of 1.0 on all four. The smaller levers scatter.

The scatter is a property of the test, not of the levers. At these loads the slowest 5 percent of requests are decided by which ones happen to collide with a long prefill, and the engine's and the simulator's slowest-5-percent sets share only 41 percent of their requests even on the baseline. The width of the pass band and the simulator's own sensitivity tell the story in one picture:

The pass band against the simulator's own sensitivity to a 2 percent changePass band, engine ± 0.15width 0.30Simulator's own range when its prefill speed moves by ±2%0.55 to 1.17, width 0.62Prefill cap, 16 sessions, seed 0, tail-mean ratio. A 2% change is well inside the simulator's calibration error.
Figure 7. A seed-matched test cannot resolve a 0.15 tolerance when a calibration-sized nudge moves the prediction by 0.62. Across seeds the two agree better: all three engine ratios for the cap at 16 sessions fall inside the simulator's range over ten seeds, and the engine's mean (0.75) equals the simulator's median. That comparison was not pre-registered, so it is reported as an observation.

On absolute tail latency the simulator is within 10 to 15 percent of the engine; the baseline at 16 sessions matched p99 to the second decimal (10.29 against 10.25 s). The conclusion for anyone validating a serving simulator against an agent trace: compare distributions over many seeds, or test only large effects.

5. Consumer GPU, server GPU

The engine tests ran on an RTX 4090. The result depends on one ratio, reload against recompute: 0.21 s against 3.66 s for a 32K context on this card, about 1 to 17. On an H100, PCIe 5.0 doubles the link bandwidth and FP8 throughput is about three times higher, so the ratio moves to about 1 to 11. The reload stays an order of magnitude cheaper than the recompute, and the policy result holds. What changes on a server GPU: the absolute latencies, the amount of DRAM that gives the same tier-to-pool ratio on an 80 GB card, and the load at which the prefill cap turns from help to harm.

Reload against recompute for a 32K context, RTX 4090 measured and H100 estimatedRTX 4090, measuredrecompute3.66 s3.66 sreload0.21 s0.21 s, one 17thH100, estimated from public specsrecomputeabout 1.2 sabout 1.2 sreloadabout 0.1 sabout 0.1 s, one 11th
Figure 8. Bars share one scale. The H100 row is an estimate from public specs, not a measurement: FP8 dense throughput about three times the 4090's, PCIe 5.0 about twice PCIe 4.0. The simulator's H100-like constant set, built from the same public figures, gives the same picture: prefetch within a few percent of the TTL policy at every load, and every policy overloaded at 128K contexts and ×4 load.
WhatCarries overWhy
Prediction-timed prefetch gains littleYesDepends on reload being far cheaper than recompute. True on both cards.
Size the host tier by its ratio to the GPU poolYes, as a ratioAn 80 GB H100 holds a larger pool, so the same 3.7× means more DRAM; the tier-to-pool ratio is the quantity that mattered.
Blocking store path costs 7 to 14% of a cold prefillMostlyProportional overhead of the copy inside the step; faster links shrink it somewhat.
Prefill cap helps with headroom, hurts when oversubscribedMechanism yes, thresholds noScheduler behaviour is vLLM's, not the card's; where the switch happens depends on pool size and traffic.
Absolute latencies (3.66 s, 0.21 s, 40 ms)NoCard-specific. Re-measure, then re-run.
Multi-GPU, NVLink, tensor parallel, MoE modelsUntestedNot in this setup. KV per token and transfer paths differ.

How this sits next to the published work

Continuum (Berkeley) keeps KV across tool calls with a TTL from reload cost and queueing delay and reports large gains on SWE-bench and BFCL. My result disagrees only in scope: on this trace the memoryfulness coefficient is negative and the reload is cheap, so the TTL is zero for 97 percent of pauses, and even forcing it positive keeps the rule within 10 percent of LRU. TokenCake offloads during function calls and partitions GPU memory by agent criticality; its finding that memory-constrained settings gain the most agrees with mine. CacheWise orders evictions by predicted reuse on its own coding traces; my data suggest the tier's size matters more than the eviction order once a tier exists. Ask the Tool (Tsinghua/Alibaba) reads progress from running tools instead of predicting from history and reports a 21 percent cut in p90 after tool calls; its first claim, that history cannot predict duration, matches the falling hazard here, and a progress-signal policy is the obvious next thing to put through this setup.

Three open proposals would consume a resume prediction: vLLM RFC #57103 (Programmable KV Cache), vLLM Router #285 (Progress-TTL) and SGLang RFC #24656 (Agent-Aware KV Cache). I am posting these results in those threads, with LRU plus an adequately sized host tier as the baseline any new policy should beat.

Limits

  • One trace source. TraceLab comes from one lab and 71 developers; tests, builds and installs are under 3 percent of calls; tool arguments are stripped, so deadlines were inferred statistically.
  • One model, one GPU. Llama 3.1 8B in FP8 on a single RTX 4090 over PCIe 4.0. Section 5 says what transfers.
  • An 8 GB host tier on this machine, so the 20 and 40 GB regimes were tested on a scaled-down workload; the full-scale 20 GB regime is simulator-only.
  • Contexts scaled so the 99th-percentile session peaks at 32K tokens. Real Claude contexts reach 1M.
  • The policy comparison is simulator-only. The engine runs cover the three levers, three seeds for two of them and one seed for the rest.

Open questions

  1. Why the native connector does not improve the tail at 32 sessions when it does at 16. Three untested candidates: it keeps GPU blocks referenced until asynchronous copies finish; it loads before scheduling; it moves 16-token blocks rather than 256-token chunks.
  2. Why LMCache's host tier served fewer reloads than modelled (70 against 184 GiB at 32 sessions). Chunk-level LRU, where an evicted early chunk makes the rest of a prefix unusable, is one candidate.
  3. Whether a memory-aware cap would keep the cap's benefit without its cost under oversubscription.

Data: TraceLab v0.0.2 (Zhu, Jacob, Ma, Pan, Wang, Krishnamurthy and Kasikci, SyFI Lab, University of Washington; arXiv 2606.30560), CC BY 4.0. Papers discussed: Continuum arXiv 2511.02230; TokenCake arXiv 2510.18586; CacheWise arXiv 2606.16824; Ask the Tool arXiv 2609.18849. Code and tables available on request.