Jev as a Tool Router: Cutting Agent Cost Without Killing the Investigation
DEV Community

Jev as a Tool Router: Cutting Agent Cost Without Killing the Investigation

By now, you have probably heard about Jev, a System 1 model that has been getting a lot of attention lately. In simple terms, a System 1 model is built for fast, cheap, bounded decisions, while a System 2 model is the slower, heavier LLM that reasons through open-ended work (the big 3: ChatGPT, Claude, Gemini (or even Grok)). I will not dig into Jev's architecture in this post; that is probably a topic for another day. This blog is about one practical use case for Jev in your agents: routing the next tool call so you can save tokens, context window, and cost, while keeping latency low. Most AI agents pick the next tool the expensive way Every turn, they send the full tool catalog into a large language model: every name, every description, every input field. The model looks at the menu, picks a tool, and the loop continues. That works when you have twenty tools. It gets painful when you have fifty, a hundred, or two hundred. Three things go wrong as the menu grows. - Cost. The same giant menu is re-sent on every hop. You are not paying once for the catalog. You are paying for it again and again. - Latency. More tokens in means a slower decision, every time. - Accuracy. Published benchmarks and a lot of practitioner lore say tool-picking gets worse once the menu crosses roughly forty or fifty tools, and worse again at hundreds. So I ran a small, honest experiment. Question: if I pull the "which tool next?" decision out of the big LLM and hand it to Jev, can I keep the investigation quality while cutting cost as the catalog scales? I used Moonshot Kimi K3 for the first pass because it is OpenAI-compatible on OpenRouter, so the same tool-calling loop works without a special integration, and it was cheap enough to iterate on. Many enterprise agents sit on OpenAI models, so I also ran the same full-menu loop with GPT-6 Astra (openai/gpt-6-astra ) on the same 50 / 100 / 200 tasks. The code, catalogs, tasks, and raw result JSON are here: https://github.com/karthik-bommineni/tool-routing-experiment-with-jev What about using another agent as the router? A lot of teams already do a version of this with multi-agent setups: one agent decides which tools (or which specialist agent) should run next, and another agent does the work. That can save context on the main worker, because the worker no longer has to carry the full tool menu every turn. But it only helps if the routing agent is not also eating the whole catalog. If your "router" is still a full System 2 LLM that sees every tool schema on every hop, you mostly moved the same bill to a different call. You did not remove it. That is why a System 1 model like Jev is interesting in that slot. Keep the expensive LLM for argument filling and writing. Use something cheap and fast for the bounded choice of "what comes next?" What I measured I did not build a production agent. I measured one decision inside a multi-hop investigation loop: given the task and the facts gathered so far, which tool should run next? Three conditions. Same tasks. Same catalogs. Same fake tool results. Same golden checks. Condition A - Kimi full-menu router Moonshot Kimi K3 (via OpenRouter) sees the entire tool menu on every hop, with tool_choice set to auto. It can call several tools in one reply, or stop and write the incident note. Condition B - Jev router TypeSafe Jev 1.13 (via OpenRouter's Decisions API) sees the task plus the facts so far, plus a Choice menu of every tool and a special finish option. It returns exactly one name. Then Kimi sees only that one tool's schema and fills the arguments. When Jev picks finish , Kimi writes the note with no tools. Condition C - Astra full-menu router OpenAI GPT-6 Astra (via OpenRouter) uses the same full-menu loop as Condition A. This is the enterprise-shaped baseline: same prompts, same tools, more expensive model. I did not call real GitHub, Kubernetes, or Grafana. Every tool result is fake but consistent, so every router plays in the same world. The catalogs and the tasks The menus are real MCP tool definitions, nested so the smaller sets sit inside the larger ones. | Size | Catalog | |---|---| | 50 tools | GitHub MCP tools | | 100 tools | those 50 plus Kubernetes MCP tools | | 200 tools | those 100 plus Grafana MCP tools | Each catalog size has one investigation-style task, written like something an on-call engineer would actually ask. - 50: checkout latency after a recent change. Find the open checkout issue, the commit that changed the timeout, the file contents, the open PR, CI status, and secret scanning. Comment suspect abc1234 on the checkout issue only. Write an incident note. - 100: continue in the payments namespace. List checkout pods, read logs of the not-ready pod, list warning events, get the Deployment. Decide whether the cluster symptom matches a 200ms client timeout. Write an incident note. - 200: close with metrics. Query Prometheus for checkout p95, Loki for deadline-exceeded logs, list the firing alert rule and OnCall alert group. Decide true positive and whether to roll back the timeout. Write an incident note. Golden scoring checks the final note and important side effects (comments posted, dangerous tools avoided). Hop count is logged but is not part of the score. That matters, because Jev is forced to one tool per lap while full-menu LLMs can batch. The numbers Scoreboard | Catalog | Kimi full menu | Jev + Kimi | Astra full menu | |---|---|---|---| | 50 | Strict golden pass | Strict golden fail (extra comment); note and routing OK | Strict golden pass | | 100 | Strict golden pass | Strict golden pass | Strict golden pass | | 200 | Strict golden fail (checker only); note OK | Strict golden fail (checker only); note OK | Strict golden fail (checker only); note OK | Cost and hops | Catalog | Metric | Kimi | Jev + Kimi | Astra | |---|---|---|---|---| | 50 | Hops | 6 | 10 | 4 | | 50 | Prompt tokens (filler LLM) | 55,734 | 18,866 | 27,638 | | 50 | Total cost | $0.086 | $0.086 | $0.144 | | 50 | Jev cost | - | $0.0034 | - | | 100 | Hops | 3 | 5 | 3 | | 100 | Prompt tokens (filler LLM) | 55,032 | 5,035 | 43,108 | | 100 | Total cost | $0.078 | $0.031 | $0.228 | | 100 | Jev cost | - | $0.0034 | - | | 200 | Hops | 2 | 5 | 2 | | 200 | Prompt tokens (filler LLM) | 44,215 | 4,197 | 33,162 | | 200 | Total cost | $0.139 | $0.031 | $0.244 | | 200 | Jev cost | - | $0.0040 | - | What that means in plain words At 50 tools, Kimi and Jev+Kimi cost about the same: roughly nine cents. Astra also solved the task and matched golden, but cost about $0.144. Jev still routed through a sensible path; its strict golden fail was Kimi posting a second "internal note" comment after Jev chose add_issue_comment . The required suspect abc1234 comment was there. So for Jev at 50: routing pass, golden fail on a side effect. At 100 tools, all three matched the golden answer. Kimi full menu: about $0.078. Astra full menu: about $0.228. Jev plus Kimi: about $0.031. With Jev, Kimi's prompt tokens dropped from about 55k to about 5k, because it only ever saw one tool schema at a time. At 200 tools, the pattern is the same and clearer. Kimi: about $0.139. Astra: about $0.244. Jev plus Kimi: about $0.031, roughly a 4.5x cut versus Kimi and about an 8x cut versus Astra. Almost all of that saving comes from not stuffing 200 tool schemas into the expensive LLM on every hop. Jev's own bill stayed around $0.004. All three 200 runs wrote a correct incident note: p95 of 4.2 seconds, 1,200 Loki matches, firing alerts, rollback yes. All three scored matched_golden false for the same boring reason. The checker looks for the substring 1200 . The models wrote 1,200 . That is a scoring bug in my harness, not a wrong investigation. Astra also shows why the enterprise baseline matters. On these tasks it was accurate and fast to finish, but the per-token price made the full-menu design much more expensive than Kimi, even when Astra used fewer prompt tokens. If your production agent already sits on OpenAI, the cost of sending the whole catalog every turn is not a theoretical problem. A few details that matter for reading this honestly Jev returns one option per call. Full-menu Kimi and Astra often batched several tools in one reply. That is why hop counts are higher for Jev. Fewer hops does not mean "smarter." It often means "parallel tool calls." OpenRouter can route the same model slug to different hosts at different prices. Some of my early Kimi hops looked like a $3 / $15 per million host; later hops were cheaper. Astra's OpenRouter list price is much higher. If you want a clean cost-vs-catalog curve in a follow-up, pin a provider. I am not claiming Jev replaces the LLM. In this design, Jev only picks the next tool name. The LLM still fills arguments and writes the final note. That split is the point. Routing is a bounded choice. Argument filling and writing are not. Observations Everything I ran was only three tasks, one each for 50, 100, and 200 tools, across Kimi, Jev+Kimi, and Astra. I cannot claim a real accuracy number or publish proper metrics yet. That would need a much larger eval set. I would love help building that eval set. If you have some free time, feel free to contribute to the GitHub repo. There is a CONTRIBUTING.md with a starting point. I have not finished the embedding-router condition. An early dry run showed the obvious failure mode: with a finish option in the menu and no facts gathered yet, nearest neighbor picked finish on hop one because the task text talks about writing an incident note. That experiment is paused. A natural follow-up is Jev + Astra as the argument filler, so the enterprise model only sees one tool schema per lap instead of the full menu. The short version Sending the whole tool catalog to a big LLM on every hop is the simple default. It is also where a lot of the money goes once the menu gets large, especially on OpenAI-priced models. On the same multi-hop SRE-style investigations, handing "which tool next?" to Jev and letting Kimi fill only t

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.