Choose an AI Agent Gateway: agentgateway Guide

How to choose an AI agent gateway by the controls it enforces, and how agentgateway meets each one in standalone and Kubernetes modes, with PoC tests.

Choose by the controls in the request path, not by model-catalog size. Write down what you must enforce on agent→model and agent→tool traffic: workload identity, MCP tool authorization, spend caps that return HTTP 429, egress and CONNECT control, secretless provider credentials, plane ownership, and exportable audit. Then require a working demo of each in your environment. I measure agentgateway against those seven controls in both of its modes, standalone and Kubernetes.

Disclosure: I work at Solo.io, which created agentgateway and sells Solo Enterprise for agentgateway. The demos I cite are public repositories, labelled open source or Solo Enterprise; every product claim links to the docs.

The architecture word is what an agentic mesh is . I am not rewriting it here.

The decision rule

Write down the controls you must enforce on agent→model and agent→tool traffic: identity, tool authz, spend, egress, credentials, plane ownership, and audit. Then require a demo of each in your environment. This page measures agentgateway, in standalone and Kubernetes modes, with the evidence linked.

Do not start from provider count or a “best LLM gateway 2026” list.

LLM proxy vs agent gateway

An LLM proxy unifies providers, routes and fails over, and attributes tokens. An agent gateway (the data plane in an agentic mesh ) still does those jobs, and also understands tool and agent protocols, binds policy to caller identity, injects provider credentials, and assumes the agent may try to CONNECT elsewhere.

If your whiteboard only has “OpenAI-compatible API + virtual keys + a log UI,” you are describing an LLM proxy, not an agent-platform hop.

Signals you are not ready for a gateway yet

One or two providers, no shared MCP plane, no multi-team agents, and spend small enough that the provider dashboard is the budget. Build the harness first. When you are ready, the smallest agentgateway footprint is the standalone binary.

Signals you need agent-aware enforcement

Multi-team agents, shared tools, an audit of which tool was called as which identity, runaway spend, or agents that can dial out. Those are platform problems. A chat-completions proxy will not fail them closed.

agentgateway’s two modes: standalone or Kubernetes

The project introduction names both modes. Same Rust proxy for HTTP, gRPC, LLM, MCP, and A2A. Different install.

  • Standalone. A single binary plus YAML. Runs on a laptop, VM, container, or edge box. No Kubernetes needed. Open-source v1.5.0 per-API-key budgets in USD or tokens are standalone-only.
  • Kubernetes. A Gateway API data plane (HTTPRoute, GRPCRoute, TCPRoute, TLSRoute) plus a controller that translates resources over xDS. Open-source budgets are token-based via global rate limiting (unit: Tokens) and return 429. Solo Enterprise for agentgateway is the commercial distribution; it adds dollar budgets on Kubernetes (EnterpriseAgentgatewayBudget).

Do not assume a standalone-only feature exists on the controller. The table below grades each control in both modes.

The scorecard

Seven rows. “MCP: yes” and “budgets: yes” are the usual false positives. Every agentgateway cell is Documented, Tested, Solo Enterprise only, or Documented (release notes).

CriterionWhat good looks likeCommon false positiveStandaloneKubernetesProof
1. IdentityPolicy binds to a cryptographic workload identity (SPIFFE or equivalent), not a header the agent can setAPI keys, virtual keys, or JWT claims the workload minted for itselfTested. Workload API, source.spiffeIdDocumented. spiffe on AgentgatewayParametersSPIFFE (OSS v1.5.0 standalone)
2. MCP tool authzPer-tool / per-server allowlists; unentitled tools hidden from tools/list“Supports MCP” or a registry screenshotDocumented. CEL mcpAuthorizationDocumented (OSS tool access ); demo on Solo Enterprise. Virtual MCP, CEL/JWTFederation (Solo Enterprise)
3. Spend caps that return 429 (dollars or tokens)Hard cap that refuses the next call with HTTP 429 before the providerA spend dashboard, an alert, or a status code you have not written into the runbookTested. OSS v1.5.0 per-key USD or tokensDocumented OSS: unit: Tokens. Dollar budgets: Solo EnterprisePer-key (OSS standalone); cost controls (Solo Enterprise)
4. Egress / CONNECTDestination authorized at CONNECT time; fail closed off-pathAn HTTP proxy the agent can walk around with NO_PROXY or /etc/hostsDocumented. CONNECT allowlist (egress proxy )Documented (release notes)Egress bypasses (pattern lab: Python CONNECT gateway in Docker Compose, not agentgateway)
5. Secretless keysAgent calls a local endpoint; gateway injects the provider credentialVault-to-env, then the process still holds sk-Tested. backendAuthDocumented. backendAuth, RFC 8693 token exchangeSecretless (OSS v1.5.0 standalone)
6. Plane ownershipYou can say who runs the request path, and point at the config“Self-hosted vs SaaS” with no mention of the binary or Gateway APITested. Binary + YAMLDocumented. Gateway API + controllerBinary , Kubernetes
7. Exportable auditMetrics and traces you can ship: who, which model or tool, why deniedA vendor UI you cannot export for a SOC 2-style reviewDocumented. Built-in OTelDocumented. Prometheus GenAI metrics, OTelSee each demo’s logs; I have not published a separate audit-export bake-off

1. Identity: SPIFFE (or equivalent workload identity)

Can policy bind to a cryptographic workload identity, not only an API key or a JWT the agent can forge?

A bearer token is a thing that can be copied. I ran the path with no certificate file on the agent, the gateway, or the model: agent-alpha got HTTP 200, agent-beta presented a valid SVID from the same trust domain and got HTTP 403 (SPIFFE identity for AI agents , OSS v1.5.0 standalone). MCP’s shipped auth still authenticates the employee more than the agent (MCP identity gap ). Do not treat “SSO for the admin UI” as the same row.

Standalone sources the mTLS identity from a SPIFFE Workload API socket and exposes source.spiffeId in CEL (v1.5.0 , CEL ). Kubernetes uses a spiffe block on AgentgatewayParameters (release notes ).

2. MCP tool authorization

Per-tool and per-server allowlists. Multi-tenant federation behind one URL. “MCP supported” is not tool authz.

Standalone MCP authorization runs CEL against tools/list and tools/call and filters disallowed tools from the list. Kubernetes tool access is the same control. I ran three customers against the same three MCP URLs and got three different tools/list results from deny-by-default CEL (multi-tenant MCP federation , Solo Enterprise, not OSS evidence). Prompt policy is not this control: AGENTS.md is not a security control .

3. Spend caps that return 429 (dollars or tokens)

Hard caps that stop the request path, not only alerts.

A budget that increments a dashboard is accounting. Open-source v1.5.0 per-key budgets (USD or tokens, Block returns 429) are standalone-only. I ran that path: a 5,000-token key refused its sixth call with a 429 and a $0.01 key refused its fourth (per-key LLM budgets ). One caveat: the gateway charges a request after the response, so the request that crosses the cap still completes and recorded spend can end above the limit; the 429 starts with the next request. Open-source Kubernetes budgets are token-based: global rate limiting with unit: Tokens returns 429. Solo Enterprise adds dollar budgets on Kubernetes; alice 429, bob on his own bucket still 200 (cost controls ). Token counts are not dollars until a catalog prices the model. I am not inventing a latency crown.

4. Egress / CONNECT control

Does the design assume all agent traffic hits the gateway, and what happens when agents CONNECT elsewhere?

The egress proxy docs describe CONNECT allowlists. In a Docker Compose pattern lab, I reproduced four bypasses, then watched all four fail behind CONNECT-time destination control (egress control ). That lab’s gateway is a Python CONNECT gateway, not agentgateway: it shows the pattern, and agentgateway’s egress proxy documents the same control. Kubernetes v1.5 release notes list the egress proxy; I have not published a Kubernetes config walkthrough. Rogue-agent Kubernetes controls is the stage map after the agent leaves the path.

5. Secretless provider keys

Agents call a local endpoint. The gateway injects the provider credential after identity and policy checks. The agent never holds the OpenAI / Anthropic / Bedrock key.

I built the split on OSS v1.5.0 standalone: inbound identity, outbound backendAuth (secretless AI agents ). Kubernetes documents RFC 8693 token exchange on the outbound path. A vault changes where the key rests, not the moment the process holds plaintext.

6. Plane ownership: standalone binary or Kubernetes Gateway API

Who owns the request path? Can you point at the process that served the request and the file or repo that configured it?

Standalone: the binary and a YAML file. Kubernetes: Gateway API resources and the controller. Solo Enterprise is the commercial distribution; label it when you use its CRDs. The introduction’s mode table is the fork: binary anywhere vs managed Kubernetes control plane.

7. Audit and observability you can export

Who called which model or tool, as which identity, and why a budget or authz denied. The introduction lists built-in metrics and tracing . On Kubernetes, the cost tracking docs cover Prometheus GenAI token metrics and llm.cost in access logs when a catalog is configured.

I have not run an observability bake-off. DeepInspect’s LLM gateway benchmarks post is useful for what to ask: policy-decision tail latency, and whether the audit write is durable. I am not publishing their numbers as mine.

How to run the evaluation

  1. Write the must-have controls into the ADR (the seven rows, or a subset you can defend).
  2. Pick the agentgateway mode that matches where agents run: standalone on a laptop or VM, or Kubernetes in your cluster.
  3. Run the PoC tests below in your environment. A lab key that never hits the hop does not test SPIFFE or tenant filter.
  4. Write the mode into the ADR: agentgateway standalone or Kubernetes, and why.
  5. Load-test policy p99 and the audit write. Be skeptical of a single median.

kagent is the runtime next door, not a gateway. Pod-versus-agent sizing is Agent Substrate .

PoC acceptance tests (copy into the ADR)

  • Identity (S, K). A SPIFFE-denied (or equivalent) agent receives 403 before any provider is contacted. A valid SVID that is not in the allowlist is not enough.
  • MCP tool authz (S, K). A tool not on the caller’s allowlist is absent from tools/list and denied on tools/call.
  • Budgets. An over-budget key receives HTTP 429 and the provider is not called. Standalone: per-key USD or tokens. Kubernetes OSS: token budget (unit: Tokens). Kubernetes dollars: Solo Enterprise.
  • Egress (S; K per release notes). An agent without the gateway path cannot reach the provider. CONNECT / hosts-file style bypasses fail.
  • Secretless (S, K). The agent environment contains no provider LLM key. A completion still returns, which means the gateway attached one.
  • Plane ownership. You can point at the process that served the request and the file or repo that configured it (standalone YAML, or Gateway API resources / Helm values). Write which mode.
  • Audit (S, K). You can export a record that names the identity, the model or tool, and the deny reason, without opening a vendor-only UI.

When not to deploy a gateway yet

Single-provider prototype, no shared tools: harness first. When you are ready, start with the agentgateway standalone binary. Move to Kubernetes mode when you standardize on Gateway API.

AGENTS.md and prompt policy are not controls. I showed that with five deleted customer records and a gateway allowlist that held (AGENTS.md is not a security control ).

Where to go next on webofmike

External primaries I would put next to the ADR, not instead of it:

The OSS project is agentgateway/agentgateway . I work at Solo.io, which created agentgateway and sells Solo Enterprise for agentgateway. This page is how to evaluate a gateway for AI agents. agentgateway is the hop, standalone and Kubernetes. I have not run a latency or QPS bake-off.

Frequently asked questions

How should a platform team choose an AI agent gateway?

Choose by the controls it enforces in the request path, not by model-catalog size. Write down what you must enforce on agent-to-model and agent-to-tool traffic: workload identity (ideally SPIFFE), MCP tool authorization, spend caps that return HTTP 429, egress and CONNECT control, secretless provider credentials, who owns the data plane, and exportable audit. Then require a working demo of each in your own environment. I measure agentgateway against those seven controls in both of its modes, standalone and Kubernetes, and link the evidence for each.

What is different about choosing a gateway for agents vs for LLM apps?

An LLM proxy mainly unifies providers, routes or fails over, and attributes tokens. An agent gateway must also understand tool and agent protocols such as MCP and A2A, bind policy to caller identity, stop spend and tool misuse mid-path, and assume agents may try to bypass the intended egress. If your scorecard has no rows for identity, tool authorization, and egress, you are still evaluating a chat-completions proxy.

Should agentgateway run standalone or on Kubernetes?

Both modes run the same agentgateway proxy for HTTP, gRPC, LLM, MCP and A2A traffic. Standalone mode is a single binary with a YAML config that runs on a laptop, VM or edge box, and in open-source v1.5.0 it is the mode with per-API-key budgets in dollars or tokens. Kubernetes mode implements Gateway API with a controller, and its open-source budgets are token-based through global rate limiting. Start standalone to prove the controls, and move to Kubernetes mode when you standardize the cluster front door on Gateway API.