I rented an A100 to test one vLLM flag
I rented an A100 for under an hour to answer one question. Does a single vLLM flag really change what a token costs? The flag was max_num_seqs : how many requests the server works on at once. I set it to 1, which is a deliberately bad setting, and then to 8. Same GPU, same model (Qwen2.5-0.5B), same prompts, eight requests in flight the whole time. The obvious way to test this is: run the old setting, run the new one, compare. I don't trust that anymore. I once watched a server I hadn't touched "get" 27% more expensive between two runs. GPUs warm up, other work comes and goes, and a single before/after quietly bakes all of that into your answer. So I interleaved it. Six runs, alternating the order: baseline → change → baseline change → baseline → change If the machine drifts during the session, the drift lands on both sides instead of making one look better. Here's what came back, in dollars per million output tokens at $1.39/hr: max_num_seqs=1 $0.748 $0.746 $0.745 max_num_seqs=8 $0.237 $0.229 $0.237 Six runs, two tight clusters, no overlap. About 68% cheaper per token, and it held every single time. Two honest caveats: - A baseline of 1 is unrealistically bad. Nobody should run production like that. The point was to prove the method, not to promise you 68%. - It's a 0.5B model. Bigger models, longer prompts and real traffic will give you a different number. That's exactly why you measure your own. What surprised me most wasn't the drop. It was how calm the numbers became once the runs were interleaved. Same server, same day, no drama. I turned this into a small open-source tool so I'd stop doing it by hand. It measures cost per million tokens with confidence intervals, and it only says CHEAPER when a change actually beats the noise: pipx install throttle-pro All six raw runs are public, if you want to check my math. Top comments (0)
Comments
No comments yet. Start the discussion.