agentgateway v1.6.0: Cost Tracking With No Catalog Config
Originally published at webofmike.com on 2026-10-06. The demo repo and every command in it were run before publishing. agentgateway v1.6.0 went GA on October 2. Two features are worth a hands-on look before anything else: a built-in LLM pricing catalog that gives you cost tracking with zero configuration, and per-key CEL rate limiting that buckets quota by any request attribute you pick. I ran both against a live gateway and a real Claude Sonnet 5 backend. Code and captured output are in themsquared/agw-16-hands-on.
Built‑in LLM pricing catalog
Cost tracking used to mean declaring rates by hand. I wrote up agentgateway's per‑API‑key budgets back in v1.5.0: a budgets list per key, a limit in tokens or USD, and a 429 before the provider ever sees the request. That post also covered the gap underneath it. Budgets only work once the gateway knows what a token costs, and in 1.5 that meant a modelCatalog block naming every model and its per‑million‑token input and output rate, maintained by hand.
v1.6.0 replaces that with a model catalog agentgateway ships and maintains itself.
Demo LLM route config
llm:
port: 4000
policies:
localRateLimit:
- type: requests
maxTokens: 3
tokensPerFill: 3
fillInterval: 60s
key: request.headers["x-api-key"]
models:
- name: claude
provider: anthropic
params:
model: claude-sonnet-5
apiKey: $ANTHROPIC_API_KEY
No modelCatalog appears anywhere.
Access log output
Because claude-sonnet-5 is a model the built‑in catalog already knows, the access log produces full cost data:
http.status=200
gen_ai.usage.input_tokens=8
gen_ai.usage.output_tokens=14
agw.ai.usage.cost.total=0.000156
cost.total=0.000156
cost.rate.input=2
cost.rate.output=10
cost.rate.input and cost.rate.output are USD per million tokens, pulled straight from the catalog. The math checks out:
(8 * 2 + 14 * 10) / 1,000,000 = 0.000156
Nothing in the config told agentgateway what this model costs; it already knew.
Getting cost fields into the access log
Those fields appear via a frontendPolicies.accessLog.add block that maps CEL expressions to log keys:
frontendPolicies:
accessLog:
add:
model.requested: llm.requestModel
model.served: llm.responseModel
tokens.input: llm.inputTokens
tokens.output: llm.outputTokens
cost.total: llm.cost.total
cost.rate.input: llm.costRates.input
cost.rate.output: llm.costRates.output
These same CEL fields are what you'd read in a budgets policy from the 1.5 post, so this isn’t a separate feature bolted on-it’s the same cost‑accounting plumbing, now backed by a catalog instead of a hand‑maintained table.
Per‑key CEL rate limiting
The second feature is localRateLimit, keyed by a CEL expression rather than a fixed value. The config above keys on request.headers["x-api-key"], with a 3‑request bucket that refills every 60 seconds. One rule, evaluated per request, buckets independently per header value:
# key A, requests 1‑3: all 200
http.status=200 ... cost.total=0.000156 ...
http.status=200 ... cost.total=0.000156 ...
http.status=200 ... cost.total=0.000156 ...
# key A, request 4, same minute
http.status=429 error="rate limit exceeded" reason=RateLimit
# key B, same minute, different header value
http.status=200 ... cost.total=0.000176
Key B never saw key A’s limit. There’s no second localRateLimit entry for it, no restart to pick up a new key. Any CEL expression over the request works as the bucket key, so this generalizes past API keys to things like a JWT claim or a source IP.
Gotcha: a 503 still spends a token
One run of the demo hit a transient upstream failure partway through:
http.status=503 error="upstream call failed: SendRequest: connection error: peer closed connection without sending TLS close_notify"
reason=UpstreamFailure
x-ratelimit-remaining still dropped on that request.
The call never reached Anthropic successfully; a 503 is the opposite of a billable response, but it still counted as one of the 3 admitted requests in the bucket. localRateLimit counts requests agentgateway admits, not requests that succeed upstream. If you're sizing maxTokens close to real traffic, budget headroom for upstream flakiness, because a bad backend day eats your quota exactly like a good one.
What this run doesn't cover
- No Kubernetes. v1.6.0's
AgentgatewayModelCRD, now on by default in the Helm chart, and K8s‑native session affinity are cluster‑side features this standalone run doesn't exercise. - One provider. The built‑in catalog covers more than Anthropic; this demo only validates the provider with a key on hand.
remoteRateLimitis a different policy, for quota shared across replicas.localRateLimitbuckets live in the single proxy instance that created them, which is exactly what makes this demo's single‑container setup representative of the behavior.
Run it yourself
Requirements: Docker and an ANTHROPIC_API_KEY. Tested on macOS (Apple silicon) against cr.agentgateway.dev/agentgateway:v1.6.0.
git clone https://github.com/themsquared/agw-16-hands-on.git
cd agw-16-hands-on
export ANTHROPIC_API_KEY=sk-ant-...
./run-demo.sh
To check the config against the v1.6.0 schema without sending any traffic:
docker run --rm -v "$PWD/config/config.yaml:/config/config.yaml:ro" \
-e ANTHROPIC_API_KEY=dummy-for-validate \
cr.agentgateway.dev/agentgateway:v1.6.0 -f /config/config.yaml --validate-only
Tear down with:
docker rm -f agw16-demo
What changed, concretely
Going from v1.5.0 to v1.6.0, the same cost‑and‑quota problem from the per‑key budgets post now needs less from you:
- No catalog to maintain.
- A quota rule that keys itself by request content instead of being written once per key.
The gotcha is the same shape either version: a limiter counts what it admits, not what succeeds, and that is worth checking against your own traffic patterns before you pick a maxTokens value.
Repo and full captured output: themsquared/agw-16-hands-on.
Frequently asked questions
Does agentgateway v1.6.0 need a model catalog configured for cost tracking?
No. v1.6.0 ships a built‑in model catalog, so a route to a known model like claude-sonnet-5 produces llm.cost.total and the input/output cost rates automatically, with no modelCatalog block anywhere in the config. Earlier versions required declaring rates by hand for every model in use.
How does agentgateway's per‑key CEL rate limiting work?
A single localRateLimit policy keyed on an expression like request.headers[…] evaluates per request and creates independent buckets for each distinct value of that expression.
Does a failed upstream request still count against an agentgateway rate limit bucket?
Yes. A request that fails upstream, such as a dropped TLS connection, still consumes a token from the bucket. localRateLimit counts requests admitted at the gateway, not successful upstream responses, so a flaky backend can eat into a tight per‑key quota without a single call completing.
Is agentgateway's remoteRateLimit the same as localRateLimit?
No. localRateLimit buckets live in memory on the single gateway instance that created them, which is exactly the single‑container setup this demo runs. remoteRateLimit is a separate policy that shares one quota across replicas, which matters once more than one gateway instance sits behind a load balancer and you need one limit enforced across all of them.
Comments
No comments yet. Start the discussion.