One field in the request made our agent 3x cheaper and 8x faster
DEV Community

One field in the request made our agent 3x cheaper and 8x faster

TL;DR. Reasoning models decide by themselves how long to think if you don't tell them. Our agent didn't - and on hard tasks the model sometimes thought for 14 minutes and 33K tokens in a single step. One field in the request body ("reasoning": {"effort": "low"} ) made a task 3x cheaper and 7-8x faster, and it solved more tasks, not fewer: 12 of 12 instead of 6 of 7. The best setting in the end was neither "always low" nor "always high" but "low, and high right after a failing test". Below: how we measured it, the tables, the code, and what it doesn't fix. About this post. The project is Altair, an open-source (Apache-2.0) AI agent for your PC and phone. I'm the author and build it largely with Claude Code. This post, the experiments and the charts were prepared by that same AI assistant; I reviewed it and stand behind it. Every number comes from our runs. This is the third post about the project: the first is about snapshots and running tests before "done", the second about cutting a browser agent's tokens by 58%. How it started We ran the agent on a cheap reasoning model (glm-5.3-flash ) and read the provider's logs. On one task the very first step took 14 minutes. The provider had time to bill it: 32,978 reasoning tokens in one step, and no answer. Our loop guard cut the stream at 160,000 characters. The cause was mundane. Reasoning models have a "how much to think" knob - reasoning_effort at OpenAI, reasoning.effort at OpenRouter, thinking in Z.ai-style APIs. If you don't send it, the provider or the model decides; OpenRouter's docs say as much. Our agent sent nothing. How we measured Easy tasks (fix a bug, answer a question about code) were solved 100% of the time under any setting, so they show no difference. We wrote six hard tasks with hidden tests - the agent never sees them, they run after it says "done": - an LRU cache with a time-to-live (11 hidden edge-case tests); - fixing a config parser from seven user bug reports; - renaming a function across a multi-file package and adding a parameter; - the 95th percentile of latencies from logs, with exclusions; - parsing "1.5M", "200K", "10 тыс"; - business days between dates with holidays, fast over a hundred years. Each task was first solved with reference code to make sure the hidden tests were fair. The agent is the real Altair, run through its CLI (altair -p ). A proxy between agent and provider logged every request: size, cached tokens, reasoning tokens, time. Cost is what OpenRouter billed. Results | Setting | Solved | Cost per task | Time per task | |---|---|---|---| | No level (before) | 6 of 7 | 12.6 m$ | 4.1 min | effort = low | 12 of 12 | 4.2 m$ | 0.55 min | effort = medium | 10 of 12 | 4.8 m$ | 0.64 min | effort = high | 12 of 12 | 6.4 m$ | 1.5 min | m$ is a thousandth of a dollar. 12.6 m$ without a level, 4.2 m$ with low : three times less. Time went from 4.1 to 0.55 minutes, 7.5x. Why "6 of 7" and not "of 12": runs without a level kept looping for 10-14 minutes, and we stopped that series early so as not to burn money. One of the six tasks (parsing "1.5M") sent the model into endless reasoning in almost every series without a level. With any explicit level - never. Where the money goes: 9,743 reasoning tokens per task on average without a level, 235 with low - 41x fewer. Output tokens cost 3.3x more than input for this model, so on hard tasks reasoning is the main bill. And on long tasks? Short tasks are 5-9 steps. We also tried 15-20-step ones: a package with six bugs in different modules, auditing twelve values across sixty files of our own code, and "read eight files in full, then answer questions about them". 9 of 9 solved in every setting, but without a level a task costs 28 m$ and takes 5.2 minutes; with an explicit level, 11-13 m$ and a little over a minute. The twist: think hard only when something broke low is cheap, high is safer. We wanted both, and tried two per-step modes: - "think about the plan" - high on the first step, thenlow ; - "think after a failure" - low , buthigh on the step right after tests failed or a tool returned an error. | Setting (18 runs each) | Solved | Cost | Time | |---|---|---|---| always low | 16 of 18 | 3.78 m$ | 0.96 min | always high | 17 of 18 | 5.72 m$ | 1.92 min | high on the first step | 15 of 18 | 4.37 m$ | 1.93 min | low , high after a failure | 17 of 18 | 3.65 m$ | 0.84 min | Careful planning didn't help - it solved the fewest. "Think after a failure" solved as many as always-high , at the cost and speed of low . It makes sense: most of an agent's work is routine (read a file, write a file, run the tests), and thinking pays off where a test just showed the first idea was wrong. One task in 18 is close to noise, honestly, but this mode is no worse than low on all three measures. It's now our default. How it looks in code The catch is that every provider has its own field: def reasoning_extra(level, base_url, model, override="auto"): """Request fields for a reasoning level ({} for "default" / unknown levels).""" if level not in LEVELS: return {} kind = dialect(base_url, model, override) if kind == "openrouter": # OpenRouter takes minimal..high; some endpoints refuse "none"/enabled=false, # so the lowest we send is "low" for "minimal" there. return {"reasoning": {"effort": "low" if level == "minimal" else level}} if kind == "zai": # Z.ai-style APIs have an on/off switch: low and minimal turn thinking off. return {"thinking": {"type": "disabled" if level in ("minimal", "low") else "enabled"}} if kind == "openai": return {"reasoning_effort": level} return {} What we stepped on along the way: - For this model OpenRouter refuses reasoning: {"enabled": false} andreasoning_effort: "none" with a 400 "Reasoning is mandatory for this endpoint". Onlyeffort: "low" works. - A Z.ai-style gateway, the other way round, takes thinking: {"type": "disabled"} and didn't answerreasoning.effort at all. - A provider that doesn't know the field may answer 400. We catch that, retry without the field and remember not to send it again. The agent loop picks the level for each step: FAILURE_RE = re.compile(r"\b\d+ failed\b|\bFAILED\b|Traceback (most recent call last)|AssertionError|" r"\bexit code [1-9]\d*\b|\bSyntaxError\b|[A-Za-z]Error:") def _reasoning_level(self) -> str | None: mode = self.settings.llm_reasoning if mode == "default": return None if mode == "adaptive": return "high" if getattr(self, "_round_failed", False) else "low" return mode After each round of tools, _round_failed is true if a tool returned an error or its output has "2 failed", a traceback or a non-zero exit code. Service calls thought more than the main ones Two more places turned up in the logs. The chat title. The agent asks the model for a title from the first message. About 95% of that answer was reasoning: ~190 tokens of "thoughts" for a five-word title. With reasoning off the answer is 12 tokens and comes 1.6-3x faster; all six sample titles were still fine. The history summary. When the context grows, the agent asks the model to condense the start of the conversation. The answer limit was 600 tokens. The model spent them on reasoning and the summary came back empty in 3 cases out of 4 (finish_reason: length ). An empty summary isn't a fold, it's a loss: the agent drops the start of the conversation instead of condensing it. With a 2,000 limit there were no empty summaries. Service calls now ask for the least reasoning, and the summary limit is 2,500. What this doesn't fix - One model. Everything was measured on glm-5.3-flash . Other models scale their levels differently, and "low" may mean something else. The principle - don't leave the level to the provider - carries over; the exact numbers don't. - Small samples. 12-18 runs per setting. A difference of one or two solved tasks is noise. A difference of 3x in cost and time is not. - Not every provider has levels. A Z.ai-style gateway only knows on/off. There "high after a failure" means "think fully", which is expensive: on our tasks "always off" on that gateway came out half the price of adaptive (5.9 vs 10.8 m$, 12 of 12 both). If cost matters most, pick "low" there. - The tasks are code. For writing, search or analytics the best level may differ. Check it yourself The core of the experiment is one agent run with a set level: # proxy.py adds the level to every agent request to the provider body = {**{"reasoning": {"effort": "low"}}, **request_body, "model": provider["model"]} # stand.py: the agent solves the task, then the hidden tests run subprocess.run([sys.executable, "altair_cli.py", "-p", "--output-format", "json", "--mode", "bypass", "--cwd", workspace, task.prompt], env=env, timeout=900) passed, detail = task.check(workspace, answer) The whole stand (the logging proxy, tasks with hidden tests and reference solutions, the summary script) and the raw results: research/agent-lab-2026-10 (python lab/facts.py recomputes every number in this post from the saved results, no keys needed). Altair's code: https://github.com/Qweezyy/AltairAgent - the level is in Settings → Agent → Reasoning level (default "Adaptive"), LLM_REASONING in .env , module pc/core/llm/reliability.py . How do you set the reasoning level in your agents - fixed, per task type, or as the work goes? Have you seen a model think far more than it needs when no level is set? Top comments (0)

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.