Rate-Limiting AI Agent Egress to Already-Allowed APIs

OpenAI agents queried Wikidata hundreds of thousands of times, a destination the allowlist permitted. The agentgateway rate limit that caps it, reproduced live.

An egress allowlist decides whether an agent can reach a destination. It has no opinion about how often. On October 7, 2026, the Wikimedia Foundation disclosed that OpenAI research agents, with no attack of any kind assigned, made unauthorized edits to Wikipedia sandbox pages starting May 12, tried and failed to exploit a hosted Etherpad instance, and sent hundreds of thousands of queries to the Wikidata Query Service (Simon Willison’s writeup , HN discussion ). The destination was never the problem. Nothing capped the rate.

Defensive security research. This post analyses publicly disclosed findings so the controls that stop them can be tested. It describes mechanisms, not procedures, and links the primary sources for anyone verifying the work.

Here’s the repo: themsquared/agent-egress-rate-limits . It’s a runnable demo of the control that does.

Why an allowlist doesn’t cover this

An allowlist is a yes/no decision about a destination. Wikidata is either on it or it isn’t, and if an agent’s job genuinely involves querying public knowledge graphs, of course it’s on it. Nothing about “permitted” says “permitted at what rate.”

I’ve written about DNS as a covert channel to a blocked destination and about a permitted registry doubling as a bidirectional message board . Both of those are about agents reaching somewhere they shouldn’t, or using a permitted channel for something other than its intended purpose. This incident is neither. Wikidata was the intended purpose. The problem is scale: a benchmark harness running unattended, a retry loop with no backoff, or a swarm of agents independently landing on the same public API will produce exactly the query volume Wikimedia described, with no malicious intent and no blocked host anywhere in the picture.

An allowlist has no answer to “how often.” Something else has to.

The control: a rate limit on the route, not the destination

agentgateway has a local rate limiting policy that attaches to any HTTP route, independent of what’s behind it. Two identical agentgateway instances, same backend, same route, with the only difference being this block:

# config/guarded.yaml
binds:
- port: 3000
  listeners:
  - protocol: HTTP
    routes:
    - policies:
        localRateLimit:
        - maxTokens: 20
          tokensPerFill: 5
          fillInterval: 1s
          type: requests
      backends:
      - host: mock-upstream-guarded:8090

maxTokens is the burst allowance. tokensPerFill and fillInterval set the steady-state refill, here 5 requests/second after the initial burst of 20 drains. type: requests matters: agentgateway also supports type: tokens for LLM-token-based limits, but this counts HTTP requests, which is the right unit for a tool call hitting an arbitrary API that has nothing to do with an LLM provider.

That last point is worth being precise about, because agentgateway already has a budget feature that sounds similar. agentgateway v1.5.0 shipped per-API-key budgets scoped to LLM token or dollar spend, enforced against the model catalog. That’s the right tool for “don’t let this key burn the month’s OpenAI bill.” It has no opinion about a tool call to Wikidata, because Wikidata isn’t an LLM provider and never touches the budget’s token-counting path. localRateLimit is a separate, generic traffic policy that doesn’t care what’s on the other end. That’s exactly what a tool-calling or MCP backend route needs.

Reproducing it

The repo stands up a mock third-party API (it answers every query and counts how many it received, so the upstream numbers are exact rather than estimated) behind two agentgateway routes: one with no policy, one with the block above.

git clone https://github.com/themsquared/agent-egress-rate-limits.git
cd agent-egress-rate-limits
docker compose up -d
bash scripts/demo.sh

A 40-request burst, fired as fast as curl can send them, through each route:

1. UNGUARDED route: no localRateLimit policy.
   Sending 40 requests as fast as curl can fire them:
  40 x 200, 0 x 429
   Upstream (the mock third-party API) saw: {"requests_received": 40, "uptime_s": 43.7}

2. GUARDED route: localRateLimit { maxTokens: 20, tokensPerFill: 5, fillInterval: 1s, type: requests }
   Same 40 requests, same speed:
  20 x 200, 20 x 429
   Upstream (the mock third-party API) saw: {"requests_received": 20, "uptime_s": 44.1}

Nothing about the destination changed between the two routes. It’s the same mock API, allowed either way. The unguarded upstream received every single request. The guarded one received exactly its burst allowance and not one more, no matter how many the client sent after that. The throttled responses carry standard rate-limit headers:

HTTP/1.1 429 Too Many Requests
x-ratelimit-limit: 20
x-ratelimit-remaining: 0
x-ratelimit-reset: 0

scripts/verify.sh asserts the claim directly rather than eyeballing the demo output: it fires 60 requests through each route and checks that the unguarded upstream received all 60, the guarded upstream received fewer, and only the guarded route ever returned a 429. On this run: unguarded received 60 of 60, guarded received 18.

What this doesn’t cover

localRateLimit is in-memory and per-instance. It doesn’t share counters across replicas or survive a restart, which is fine for one gateway but not for an exact global count across a fleet. agentgateway also supports remote rate limiting backed by a shared store over the Envoy rate-limit gRPC protocol for that case, out of scope for this repo.

It also doesn’t tell you why the volume happened. It caps it regardless of cause, retry storm, benchmark harness, or swarm, which is the actual point: you don’t have to classify intent correctly to bound the blast radius. And the mock upstream is not Wikidata. It’s a Python HTTP server that counts requests. The burst sizes here (tens of requests) are nowhere near “hundreds of thousands of queries”; they’re sized so the policy’s effect is visible in a few seconds on a laptop. The mechanism, a token bucket in front of an allowed destination, is what scales, not these specific numbers.

No agent framework appears anywhere in this repo. curl in a loop stands in for an agent’s tool-calling loop, because the gateway doesn’t know or care what’s on the client side of the connection. That’s exactly why the control works regardless of which model, framework, or orchestration layer is making the calls.

Where this fits

The allowlist decides if. DNS filtering and covert-channel detection watch for an allowed channel being misused for something it wasn’t meant for. A rate limit on the route answers the question none of those cover: an agent is allowed to be here, and it’s asking a lot faster than anyone intended. The fix for the Wikidata case wasn’t a smarter allowlist. It was a limit on how often “allowed” gets to mean “unlimited.”

Repo: themsquared/agent-egress-rate-limits . I work at Solo.io, which makes agentgateway.

Frequently asked questions

Does an egress allowlist stop an AI agent from overwhelming an allowed API?

No. An allowlist answers one question: is this agent permitted to reach this host? It says nothing about how often. A destination can be entirely legitimate, like a public API, and still take damage from an agent stuck in a retry loop or a swarm of agents independently deciding the same endpoint is useful.

How do I rate limit AI agent traffic to a third-party API with agentgateway?

Attach a localRateLimit policy to the route in front of that backend: maxTokens sets the burst size, tokensPerFill and fillInterval set the refill rate, and type: requests counts requests rather than LLM tokens. Requests over the limit get an HTTP 429 from the gateway before they ever reach the upstream.

Is agentgateway's rate limiting specific to LLM provider calls?

No. agentgateway also has per-API-key budgets scoped to LLM token and dollar spend, but localRateLimit is a separate, generic traffic policy. It attaches to any HTTP route and backend, which is what makes it the right control for a tool call or MCP backend hitting an arbitrary third-party API, not just an LLM provider.

What did OpenAI's agents do to Wikimedia's infrastructure?

The Wikimedia Foundation disclosed on October 7, 2026 that OpenAI research agents made unauthorized edits to Wikipedia sandbox pages starting May 12, attempted to exploit a hosted Etherpad instance, and sent hundreds of thousands of queries to the Wikidata Query Service, all during ordinary research tasks with no attack assigned.