# Rate-Limiting AI Agent Egress to Already-Allowed APIs

> OpenAI agents queried Wikidata hundreds of thousands of times, a destination the allowlist permitted. The agentgateway rate limit that caps it, reproduced live.

- Canonical URL: https://webofmike.com/agent-egress-rate-limits/
- Author: Mike Moore (https://webofmike.com/about/)
- Published: 2026-10-08
- Last modified: 2026-10-08
- Tags: AI Agents, Security, AI Gateways, Platform Engineering
- Cite as: Mike Moore, "Rate-Limiting AI Agent Egress to Already-Allowed APIs", Web of Mike (webofmike.com), 2026-10-08. https://webofmike.com/agent-egress-rate-limits/


An egress allowlist decides whether an agent can reach a destination. It has no opinion about how often. On October 7, 2026, the Wikimedia Foundation disclosed that OpenAI research agents, with no attack of any kind assigned, made unauthorized edits to Wikipedia sandbox pages starting May 12, tried and failed to exploit a hosted Etherpad instance, and sent hundreds of thousands of queries to the Wikidata Query Service ([Simon Willison's writeup](https://simonwillison.net/2026/Oct/7/openai-rogue-agents-wikimedia/), [HN discussion](https://news.ycombinator.com/item?id=49968105)). The destination was never the problem. Nothing capped the rate.

Here's the repo: [themsquared/agent-egress-rate-limits](https://github.com/themsquared/agent-egress-rate-limits). It's a runnable demo of the control that does.

## Why an allowlist doesn't cover this

An allowlist is a yes/no decision about a destination. Wikidata is either on it or it isn't, and if an agent's job genuinely involves querying public knowledge graphs, of course it's on it. Nothing about "permitted" says "permitted at what rate."

I've [written about DNS as a covert channel to a blocked destination](/agent-dns-egress-covert-channel/) and about a [permitted registry doubling as a bidirectional message board](/registry-covert-channel/). Both of those are about agents reaching somewhere they shouldn't, or using a permitted channel for something other than its intended purpose. This incident is neither. Wikidata was the intended purpose. The problem is scale: a benchmark harness running unattended, a retry loop with no backoff, or a swarm of agents independently landing on the same public API will produce exactly the query volume Wikimedia described, with no malicious intent and no blocked host anywhere in the picture.

An allowlist has no answer to "how often." Something else has to.

## The control: a rate limit on the route, not the destination

agentgateway has a [local rate limiting policy](https://agentgateway.dev/docs/standalone/main/documentation/configuration/resiliency/rate-limits/) that attaches to any HTTP route, independent of what's behind it. Two identical agentgateway instances, same backend, same route, with the only difference being this block:

```yaml
# config/guarded.yaml
binds:
- port: 3000
  listeners:
  - protocol: HTTP
    routes:
    - policies:
        localRateLimit:
        - maxTokens: 20
          tokensPerFill: 5
          fillInterval: 1s
          type: requests
      backends:
      - host: mock-upstream-guarded:8090
```

`maxTokens` is the burst allowance. `tokensPerFill` and `fillInterval` set the steady-state refill, here 5 requests/second after the initial burst of 20 drains. `type: requests` matters: agentgateway also supports `type: tokens` for LLM-token-based limits, but this counts HTTP requests, which is the right unit for a tool call hitting an arbitrary API that has nothing to do with an LLM provider.

That last point is worth being precise about, because agentgateway already has a budget feature that sounds similar. [agentgateway v1.5.0 shipped per-API-key budgets](/agentgateway-per-key-llm-budgets/) scoped to LLM token or dollar spend, enforced against the model catalog. That's the right tool for "don't let this key burn the month's OpenAI bill." It has no opinion about a tool call to Wikidata, because Wikidata isn't an LLM provider and never touches the budget's token-counting path. `localRateLimit` is a separate, generic traffic policy that doesn't care what's on the other end. That's exactly what a tool-calling or MCP backend route needs.

## Reproducing it

The repo stands up a mock third-party API (it answers every query and counts how many it received, so the upstream numbers are exact rather than estimated) behind two agentgateway routes: one with no policy, one with the block above.

```bash
git clone https://github.com/themsquared/agent-egress-rate-limits.git
cd agent-egress-rate-limits
docker compose up -d
bash scripts/demo.sh
```

A 40-request burst, fired as fast as curl can send them, through each route:

```
1. UNGUARDED route: no localRateLimit policy.
   Sending 40 requests as fast as curl can fire them:
  40 x 200, 0 x 429
   Upstream (the mock third-party API) saw: {"requests_received": 40, "uptime_s": 43.7}

2. GUARDED route: localRateLimit { maxTokens: 20, tokensPerFill: 5, fillInterval: 1s, type: requests }
   Same 40 requests, same speed:
  20 x 200, 20 x 429
   Upstream (the mock third-party API) saw: {"requests_received": 20, "uptime_s": 44.1}
```

Nothing about the destination changed between the two routes. It's the same mock API, allowed either way. The unguarded upstream received every single request. The guarded one received exactly its burst allowance and not one more, no matter how many the client sent after that. The throttled responses carry standard rate-limit headers:

```
HTTP/1.1 429 Too Many Requests
x-ratelimit-limit: 20
x-ratelimit-remaining: 0
x-ratelimit-reset: 0
```

`scripts/verify.sh` asserts the claim directly rather than eyeballing the demo output: it fires 60 requests through each route and checks that the unguarded upstream received all 60, the guarded upstream received fewer, and only the guarded route ever returned a 429. On this run: unguarded received 60 of 60, guarded received 18.

## What this doesn't cover

`localRateLimit` is in-memory and per-instance. It doesn't share counters across replicas or survive a restart, which is fine for one gateway but not for an exact global count across a fleet. agentgateway also supports remote rate limiting backed by a shared store over the Envoy rate-limit gRPC protocol for that case, out of scope for this repo.

It also doesn't tell you why the volume happened. It caps it regardless of cause, retry storm, benchmark harness, or swarm, which is the actual point: you don't have to classify intent correctly to bound the blast radius. And the mock upstream is not Wikidata. It's a Python HTTP server that counts requests. The burst sizes here (tens of requests) are nowhere near "hundreds of thousands of queries"; they're sized so the policy's effect is visible in a few seconds on a laptop. The mechanism, a token bucket in front of an allowed destination, is what scales, not these specific numbers.

No agent framework appears anywhere in this repo. `curl` in a loop stands in for an agent's tool-calling loop, because the gateway doesn't know or care what's on the client side of the connection. That's exactly why the control works regardless of which model, framework, or orchestration layer is making the calls.

## Where this fits

The allowlist decides if. DNS filtering and covert-channel detection watch for an allowed channel being misused for something it wasn't meant for. A rate limit on the route answers the question none of those cover: an agent is allowed to be here, and it's asking a lot faster than anyone intended. The fix for the Wikidata case wasn't a smarter allowlist. It was a limit on how often "allowed" gets to mean "unlimited."

Repo: [themsquared/agent-egress-rate-limits](https://github.com/themsquared/agent-egress-rate-limits). I work at Solo.io, which makes agentgateway.

