When a coding agent calls a tool, its KV cache sits in GPU memory until the agent comes back. Several research groups and open RFCs in vLLM and SGLang propose to predict the return and move the cache out and back in around it. I tested that on 744,000 real tool calls and on a real vLLM server. The prediction does not pay. What moves the tail is the size of the host memory tier, the store path, and one scheduler flag.
The question
A coding agent spends much of its wall-clock time waiting on tools: test runs, builds, package installs. While it waits, its KV cache occupies the scarcest resource on an inference server, GPU memory. The open proposals are built on a simple idea: when a tool call starts, move the cache to host memory; shortly before the agent returns, bring it back; time the return by a prediction from past tool durations.
I had a supply-chain version of the same idea. Tool durations might cluster around a typical length, the way deliveries do, so the probability of return would rise toward that length and a serving system could learn, per tool type, when to demote and when to prefetch. Two questions followed. Can the prediction be good enough to pay off? And if it cannot, what actually decides how long an agent waits for its next token after a tool call?
What I did
- Measured the pause distributions on 743,819 real tool calls from Claude Code and Codex. The data is TraceLab, released by the University of Washington under CC BY 4.0: 8,058 sessions from 71 developers, with tool names and timings but no arguments.
- Measured the cost of every operation on one RTX 4090 running vLLM 0.30 with Llama 3.1 8B in FP8: prefill, decode, store to host memory, reload.
- Compared retention and prefetch policies in a discrete-event simulator built from those two inputs: 636 runs over load, host tier size, context length and hardware constants.
- Checked the simulator against the real engine, fourteen one-hour runs, and used those runs to test what moves the tail.
Every number below comes from a script or a table in the project.
1. Tool pauses have a falling hazard
Most pauses are short and the long ones hold the time. 70 percent of Claude Code pauses end within a second; the 4.9 percent of tool calls over a minute hold 92 percent of all tool time. Human approval waits are 37 percent of Claude's tool time on their own.
The shape of the long pauses decides whether prefetch can be timed. The idea I started with needs a return rate that rises toward a typical duration. The data shows the opposite. Pick how long the agent has already waited:
The only sharp structure in the data is harness deadlines. Codex polls a running command at its yield deadline of 1, 5, 10 or 30 seconds, so a three-minute test reaches the engine as a chain of short pauses separated by small LLM steps; 30 percent of Codex pauses end on such a deadline. Claude Code pauses spike at its 120 and 600 second Bash timeouts and at sleeps the agent writes itself. A gateway that parses the tool call sees these values directly.
Better keys do improve the forecast. Conditioning the remaining-time distribution on tool class, command skeleton and developer, with shrinkage toward the parent level when a key has few observations, cuts the upper-quantile error for Claude by 27 to 37 percent, and the shrinkage works where Continuum's rule of a hundred observations per tool stays at the global baseline. At the low quantiles that a prefetch rule uses when stalls are expensive, the key barely matters. A better forecast never turns into a serving gain, for the reason in the next section.
2. Prediction-timed retention and prefetch do not beat LRU plus a host tier
The simulator replays TraceLab sessions closed-loop against a model of continuous batching, an HBM token pool, an LRU host tier with write-through stores and one shared PCIe link, calibrated on twelve single-session measurements. Contexts are scaled so the 99th-percentile session peaks at 32K tokens; load is counted in multiples of the eight sessions whose KV just fits the GPU pool. Six policies ran on it:
- P0, no host tier at all.
- P1, LRU with a host tier, which is what vLLM plus LMCache does out of the box.
- P2, Continuum's time-to-live rule, implemented from the paper's equation with per-tool duration distributions.
- P4-D, a deadline-aware prefetch that reloads just before a known harness deadline. Deadlines are exact here, which makes it an upper bound.
- P5, an oracle that knows every return time.
- P5-B, a strict Belady oracle: it never pins, evicts the session that returns latest, and prefetches each session just before its return.
The pre-registered success rule was a gain of more than 10 percent on p99 over the TTL policy, on all traffic or on the deadline subset, or a 10 percent gain in sessions served within a one-second SLO. Over 32 stable configurations:
- Continuum's TTL rule pins the cache for 3.4 percent of tool pauses on this trace and behaves like LRU. The trace's memoryfulness coefficient is negative (−0.35), a case the Continuum paper notes as possible but did not observe on its own workloads.
- Deadline prefetch clears 10 percent in 6 configurations, never in every seed. On the deadline subset, 9.
- The Belady oracle clears 10 percent in 9, 5 of them in every seed, most of them at the lightest load. Its median p99 ratio to the TTL policy is 0.96.
- The goodput rule never passes.
- Forcing the memoryfulness coefficient positive makes the TTL pin up to two thirds of tool pauses. The TTL policy stays within about 10 percent of LRU at every setting.
The reason is Figure 1. A pin can save at most one reload plus the queueing it avoids, and with a host tier present that rarely outweighs the GPU capacity the pin holds. The PCIe link runs at 1 to 3 percent utilisation; its queueing delay comes from large stores of cold prompts, not from bursts of reloads. The negative result covers KV placement on a PCIe-attached host tier with that tier present. Remote or cross-node tiers, setups with no host tier, and uses of prediction for scheduling or admission were not tested.
3. What moves the tail
Engine evidence first. Every number in this section comes from a one-hour replay of real TraceLab sessions against vLLM 0.30 with LMCache on the 4090, with prompts of the right lengths and session-consistent prefixes so caching and eviction behave as on real traffic.
The host tier, sized well above the GPU pool
This machine cannot pin more than 8 GB without swapping, so I scaled the workload down five times and let 4 and 8 GB stand for host tiers at 1.9 and 3.7 times the GPU KV pool, the ratios that 20 and 40 GB would have at full scale. Doubling the tier cut the tail mean of post-tool TTFT by 71, 39 and 29 percent on three seeds at 16 sessions, and by 44 percent at 32.
What the slowest requests are made of follows that ratio. Each square below is one of the 100 slowest requests in a run:
The store path
LMCache's in-process connector writes the KV through to host memory inside the engine's step, so each forward pass waits for the copy. vLLM's own offloading connector (--kv-offloading-backend native) copies on a separate CUDA stream.
| Batch 1, median of 5 | 1K tokens | 8K | 32K |
|---|---|---|---|
| Cold prefill, no connector | 71 ms | 571 ms | 3,648 ms |
| Cold prefill with LMCache store | 77 ms (+7%) | 649 ms (+14%) | 4,019 ms (+10%) |
| Cold prefill with native store | 71 ms (0%) | 569 ms (0%) | 3,670 ms (+1%) |
| Reload from host memory, LMCache | 21 ms | 63 ms | 210 ms |
| Reload from host memory, native | 36 ms | 125 ms | 409 ms |
Under load the blocking store adds up. At 32 sessions LMCache stored 2.7 TiB during the hour and held the GPU for 238 seconds of it, 6.6 percent. Native copied a similar volume overlapped with compute. On the engine, native cut p95 by 41 percent at 16 sessions; at 32 sessions it halved the median and left the tail where it was. Its reloads run at about half LMCache's speed and did not show up in the tail at either load. On vLLM 0.30 the native connector is the safer default. Why the tail at 32 sessions does not improve is an open question: reload speed and store cost are both ruled out by the simulator.
The prefill chunk cap
long_prefill_token_threshold caps how many tokens one request may prefill per engine step, so short prefills queued behind a long one still get budget. At 16 sessions it improved the tail mean on all three seeds, by 17 to 35 percent, with p95 mixed. At 32 sessions the GPU pool is oversubscribed, 37 percent of post-tool requests recompute their context, and the cap backfires: p95 up 13 percent, the median 2.3 times worse. With the cap on, the engine keeps more prefills in flight at once, GPU KV usage rises from 46 to 53 percent, and fewer requests are served from the GPU cache. Set it last, and only with memory headroom.
What an operator should do, in order
| Order | Lever | Decision | Engine evidence |
|---|---|---|---|
| 1 | Host tier size | Size it well above the GPU KV pool: 3.7× rather than 1.9×. | Tail mean −71, −39, −29% on three seeds at 16 sessions; −44% at 32. |
| 2 | Store path | On vLLM 0.30, use the native offloading connector rather than LMCache's in-process connector. | p95 −41% at 16 sessions; median halved at 32; store penalty +1% against +10% on a cold 32K prefill. |
| 3 | Prefill chunk cap | Set long_prefill_token_threshold only when the GPU pool has headroom. | Tail mean −17 to −35% at 16 sessions; p95 +13% and median ×2.3 when oversubscribed. |
| 4 | Retention timing, prediction-timed prefetch | Leave to LRU once a host tier exists. | Nothing on the engine. In simulation, within seed noise. |
4. What the simulator can and cannot predict
Before any lever result came in, I fixed a pass rule: for a lever at a given load and seed, the engine's and the simulator's ratios (lever on over off) must agree within 0.15 on p95 and on the tail mean, absolute values within 15 percent or 30 ms, and direction must match when the predicted effect exceeds 15 percent. Three of ten checks passed.
The scatter is a property of the test, not of the levers. At these loads the slowest 5 percent of requests are decided by which ones happen to collide with a long prefill, and the engine's and the simulator's slowest-5-percent sets share only 41 percent of their requests even on the baseline. The width of the pass band and the simulator's own sensitivity tell the story in one picture:
On absolute tail latency the simulator is within 10 to 15 percent of the engine; the baseline at 16 sessions matched p99 to the second decimal (10.29 against 10.25 s). The conclusion for anyone validating a serving simulator against an agent trace: compare distributions over many seeds, or test only large effects.
5. Consumer GPU, server GPU
The engine tests ran on an RTX 4090. The result depends on one ratio, reload against recompute: 0.21 s against 3.66 s for a 32K context on this card, about 1 to 17. On an H100, PCIe 5.0 doubles the link bandwidth and FP8 throughput is about three times higher, so the ratio moves to about 1 to 11. The reload stays an order of magnitude cheaper than the recompute, and the policy result holds. What changes on a server GPU: the absolute latencies, the amount of DRAM that gives the same tier-to-pool ratio on an 80 GB card, and the load at which the prefill cap turns from help to harm.
| What | Carries over | Why |
|---|---|---|
| Prediction-timed prefetch gains little | Yes | Depends on reload being far cheaper than recompute. True on both cards. |
| Size the host tier by its ratio to the GPU pool | Yes, as a ratio | An 80 GB H100 holds a larger pool, so the same 3.7× means more DRAM; the tier-to-pool ratio is the quantity that mattered. |
| Blocking store path costs 7 to 14% of a cold prefill | Mostly | Proportional overhead of the copy inside the step; faster links shrink it somewhat. |
| Prefill cap helps with headroom, hurts when oversubscribed | Mechanism yes, thresholds no | Scheduler behaviour is vLLM's, not the card's; where the switch happens depends on pool size and traffic. |
| Absolute latencies (3.66 s, 0.21 s, 40 ms) | No | Card-specific. Re-measure, then re-run. |
| Multi-GPU, NVLink, tensor parallel, MoE models | Untested | Not in this setup. KV per token and transfer paths differ. |
How this sits next to the published work
Continuum (Berkeley) keeps KV across tool calls with a TTL from reload cost and queueing delay and reports large gains on SWE-bench and BFCL. My result disagrees only in scope: on this trace the memoryfulness coefficient is negative and the reload is cheap, so the TTL is zero for 97 percent of pauses, and even forcing it positive keeps the rule within 10 percent of LRU. TokenCake offloads during function calls and partitions GPU memory by agent criticality; its finding that memory-constrained settings gain the most agrees with mine. CacheWise orders evictions by predicted reuse on its own coding traces; my data suggest the tier's size matters more than the eviction order once a tier exists. Ask the Tool (Tsinghua/Alibaba) reads progress from running tools instead of predicting from history and reports a 21 percent cut in p90 after tool calls; its first claim, that history cannot predict duration, matches the falling hazard here, and a progress-signal policy is the obvious next thing to put through this setup.
Three open proposals would consume a resume prediction: vLLM RFC #57103 (Programmable KV Cache), vLLM Router #285 (Progress-TTL) and SGLang RFC #24656 (Agent-Aware KV Cache). I am posting these results in those threads, with LRU plus an adequately sized host tier as the baseline any new policy should beat.
Limits
- One trace source. TraceLab comes from one lab and 71 developers; tests, builds and installs are under 3 percent of calls; tool arguments are stripped, so deadlines were inferred statistically.
- One model, one GPU. Llama 3.1 8B in FP8 on a single RTX 4090 over PCIe 4.0. Section 5 says what transfers.
- An 8 GB host tier on this machine, so the 20 and 40 GB regimes were tested on a scaled-down workload; the full-scale 20 GB regime is simulator-only.
- Contexts scaled so the 99th-percentile session peaks at 32K tokens. Real Claude contexts reach 1M.
- The policy comparison is simulator-only. The engine runs cover the three levers, three seeds for two of them and one seed for the rest.
Open questions
- Why the native connector does not improve the tail at 32 sessions when it does at 16. Three untested candidates: it keeps GPU blocks referenced until asynchronous copies finish; it loads before scheduling; it moves 16-token blocks rather than 256-token chunks.
- Why LMCache's host tier served fewer reloads than modelled (70 against 184 GiB at 32 sessions). Chunk-level LRU, where an evicted early chunk makes the rest of a prefix unusable, is one candidate.
- Whether a memory-aware cap would keep the cap's benefit without its cost under oversubscription.
Data: TraceLab v0.0.2 (Zhu, Jacob, Ma, Pan, Wang, Krishnamurthy and Kasikci, SyFI Lab, University of Washington; arXiv 2606.30560), CC BY 4.0. Papers discussed: Continuum arXiv 2511.02230; TokenCake arXiv 2510.18586; CacheWise arXiv 2606.16824; Ask the Tool arXiv 2609.18849. Code and tables available on request.