DEV Community

agentgateway v1.6.0: Cost Tracking With No Catalog Config

Originally published at webofmike.com on 2026-10-06. The demo repo and every command in it were run before publishing. agentgateway v1.6.0 went GA on October 2. Two features are worth a hands-on look before anything else: a built-in LLM pricing catalog that gives you cost tracking with zero configuration, and per-key CEL rate limiting that buckets quota by any request attribute you pick. I ran both against a live gateway and a real Claude Sonnet 5 backend. Code and captured output are in themsquared/agw-16-hands-on.

Built‑in LLM pricing catalog

Cost tracking used to mean declaring rates by hand. I wrote up agentgateway's per‑API‑key budgets back in v1.5.0: a budgets list per key, a limit in tokens or USD, and a 429 before the provider ever sees the request. That post also covered the gap underneath it. Budgets only work once the gateway knows what a token costs, and in 1.5 that meant a modelCatalog block naming every model and its per‑million‑token input and output rate, maintained by hand.

v1.6.0 replaces that with a model catalog agentgateway ships and maintains itself.

Demo LLM route config

llm:
  port: 4000
  policies:
    localRateLimit:
      - type: requests
        maxTokens: 3
        tokensPerFill: 3
        fillInterval: 60s
        key: request.headers["x-api-key"]
    models:
      - name: claude
        provider: anthropic
        params:
          model: claude-sonnet-5
        apiKey: $ANTHROPIC_API_KEY

No modelCatalog appears anywhere.

Access log output

Because claude-sonnet-5 is a model the built‑in catalog already knows, the access log produces full cost data:

http.status=200
gen_ai.usage.input_tokens=8
gen_ai.usage.output_tokens=14
agw.ai.usage.cost.total=0.000156
cost.total=0.000156
cost.rate.input=2
cost.rate.output=10

cost.rate.input and cost.rate.output are USD per million tokens, pulled straight from the catalog. The math checks out:

(8 * 2 + 14 * 10) / 1,000,000 = 0.000156

Nothing in the config told agentgateway what this model costs; it already knew.

Getting cost fields into the access log

Those fields appear via a frontendPolicies.accessLog.add block that maps CEL expressions to log keys:

frontendPolicies:
  accessLog:
    add:
      model.requested: llm.requestModel
      model.served: llm.responseModel
      tokens.input: llm.inputTokens
      tokens.output: llm.outputTokens
      cost.total: llm.cost.total
      cost.rate.input: llm.costRates.input
      cost.rate.output: llm.costRates.output

These same CEL fields are what you'd read in a budgets policy from the 1.5 post, so this isn’t a separate feature bolted on-it’s the same cost‑accounting plumbing, now backed by a catalog instead of a hand‑maintained table.

Per‑key CEL rate limiting

The second feature is localRateLimit, keyed by a CEL expression rather than a fixed value. The config above keys on request.headers["x-api-key"], with a 3‑request bucket that refills every 60 seconds. One rule, evaluated per request, buckets independently per header value:

# key A, requests 1‑3: all 200
http.status=200 ... cost.total=0.000156 ...
http.status=200 ... cost.total=0.000156 ...
http.status=200 ... cost.total=0.000156 ...
# key A, request 4, same minute
http.status=429 error="rate limit exceeded" reason=RateLimit
# key B, same minute, different header value
http.status=200 ... cost.total=0.000176

Key B never saw key A’s limit. There’s no second localRateLimit entry for it, no restart to pick up a new key. Any CEL expression over the request works as the bucket key, so this generalizes past API keys to things like a JWT claim or a source IP.

Gotcha: a 503 still spends a token

One run of the demo hit a transient upstream failure partway through:

http.status=503 error="upstream call failed: SendRequest: connection error: peer closed connection without sending TLS close_notify"
reason=UpstreamFailure
x-ratelimit-remaining still dropped on that request.

The call never reached Anthropic successfully; a 503 is the opposite of a billable response, but it still counted as one of the 3 admitted requests in the bucket. localRateLimit counts requests agentgateway admits, not requests that succeed upstream. If you're sizing maxTokens close to real traffic, budget headroom for upstream flakiness, because a bad backend day eats your quota exactly like a good one.

What this run doesn't cover

  • No Kubernetes. v1.6.0's AgentgatewayModel CRD, now on by default in the Helm chart, and K8s‑native session affinity are cluster‑side features this standalone run doesn't exercise.
  • One provider. The built‑in catalog covers more than Anthropic; this demo only validates the provider with a key on hand.
  • remoteRateLimit is a different policy, for quota shared across replicas. localRateLimit buckets live in the single proxy instance that created them, which is exactly what makes this demo's single‑container setup representative of the behavior.

Run it yourself

Requirements: Docker and an ANTHROPIC_API_KEY. Tested on macOS (Apple silicon) against cr.agentgateway.dev/agentgateway:v1.6.0.

git clone https://github.com/themsquared/agw-16-hands-on.git
cd agw-16-hands-on
export ANTHROPIC_API_KEY=sk-ant-...
./run-demo.sh

To check the config against the v1.6.0 schema without sending any traffic:

docker run --rm -v "$PWD/config/config.yaml:/config/config.yaml:ro" \
  -e ANTHROPIC_API_KEY=dummy-for-validate \
  cr.agentgateway.dev/agentgateway:v1.6.0 -f /config/config.yaml --validate-only

Tear down with:

docker rm -f agw16-demo

What changed, concretely

Going from v1.5.0 to v1.6.0, the same cost‑and‑quota problem from the per‑key budgets post now needs less from you:

  • No catalog to maintain.
  • A quota rule that keys itself by request content instead of being written once per key.

The gotcha is the same shape either version: a limiter counts what it admits, not what succeeds, and that is worth checking against your own traffic patterns before you pick a maxTokens value.

Repo and full captured output: themsquared/agw-16-hands-on.

Frequently asked questions

Does agentgateway v1.6.0 need a model catalog configured for cost tracking?
No. v1.6.0 ships a built‑in model catalog, so a route to a known model like claude-sonnet-5 produces llm.cost.total and the input/output cost rates automatically, with no modelCatalog block anywhere in the config. Earlier versions required declaring rates by hand for every model in use.

How does agentgateway's per‑key CEL rate limiting work?
A single localRateLimit policy keyed on an expression like request.headers[…] evaluates per request and creates independent buckets for each distinct value of that expression.

Does a failed upstream request still count against an agentgateway rate limit bucket?
Yes. A request that fails upstream, such as a dropped TLS connection, still consumes a token from the bucket. localRateLimit counts requests admitted at the gateway, not successful upstream responses, so a flaky backend can eat into a tight per‑key quota without a single call completing.

Is agentgateway's remoteRateLimit the same as localRateLimit?
No. localRateLimit buckets live in memory on the single gateway instance that created them, which is exactly the single‑container setup this demo runs. remoteRateLimit is a separate policy that shares one quota across replicas, which matters once more than one gateway instance sits behind a load balancer and you need one limit enforced across all of them.

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.