You make GPU decisions on one benchmark run. I ran the same one 14 times.
Nearly every serving benchmark number, published or internal, comes from running the thing once. I ran one identical vLLM benchmark 14 times: throughput reproduces to 0.02%, the latency tail swings 18%, and one config turned out to be bistable rather than noisy.