AI Agent Security on Kubernetes: Which Controls Actually Held

A month of AI agent security labs on identity, egress, audit logs, MCP tool poisoning, runtimes and evals. The controls that held all sat outside the agent.

I spent the last month building small, runnable labs for the ways AI agents break the infrastructure they run on: credentials, egress, audit logs, MCP tool definitions, runtimes and evals. Read together, they make one argument. Every control that failed read a field the agent chooses or lived inside the agent’s own process. Every control that held sat at a chokepoint the agent cannot reach and decided on something the agent cannot set. This post pulls the first-party findings from those labs into one place, links the full write-ups, and is honest about the part that did not work.

Defensive security research. This post analyses publicly disclosed findings and reproduces them in an isolated local lab — Docker Compose or a single Python process, no live systems and no third-party infrastructure — so that the controls that stop them can be tested. It describes mechanisms, not procedures, and links the primary sources for anyone verifying the work.

Every number below comes from one of the linked posts or repos. Where a post’s claim was checked on a specific date, I say so, because a few of these things are still moving.

The rule that held across all six areas

Here is the pattern, one row per lab:

What failedThe field the agent controlledWhat held
Proxy rule: writes are blockedThe HTTP method (a write sent as a GET)Gateway that authorizes the resolved destination at CONNECT
Hostname denylistThe name in the URI (an address literal, a relay)Same gateway, which resolves names itself
NO_PROXY exemption/etc/hosts and its own environmentNo route off the network except through the gateway
AGENTS.md rule and in-band classifierThe agent’s context windowmcpAuthorization at agentgateway
MCP tool name allowlistThe tool descriptionPinning name, description and schema
Transcript monitorThe transcript itselfA proxy log the agent has no mount for
RBAC on Secret objectsA key already in the pod’s environmentNo provider key in the agent at all

None of the failed controls had a bug. Each one did exactly what it was built to do, against an input the agent was free to change. A rule written against a field the workload picks is a request, and an agent working toward a task will treat it as one.

Identity: the credential you give the agent is the credential an attacker gets

The cleanest demonstration was the dullest one. In Your Agent’s LLM Key Survives Every Kubernetes Secret Control I turned on the restricted Pod Security Standard, gave the agent’s service account no verbs on secrets, set automountServiceAccountToken: false, dropped every capability and made the root filesystem read-only. All of it verified. Then one file read inside the container returned LLM_API_KEY, and a third pod with no permissions replayed it for a 200.

The reason is structural. RBAC authorizes requests to the API server. A projected Secret never makes one from the agent’s side, because the kubelet already wrote the value into the environment at pod start. get secrets: no is true about a request the agent was never going to make.

So the fix is not a better secret store. In Your AI Agent Should Not Hold the LLM API Key the agent holds a 60-second identity token and nothing else, and agentgateway attaches the real provider credential on the way out. The LiteLLM PyPI compromise that motivated it harvested environment variables, .env files and Kubernetes configs, which is a precise list of where an agent’s key usually lives.

Scoping matters as much as holding. kagent v0.10.0 shipped audience-bound agent tokens through KAGENT_STS_RESOURCE and KAGENT_STS_AUDIENCE, and both default to empty. Unset, the exchange sends no scoping parameters at all, and that failure is silent. And the protocol layer is not going to hand you this yet: as of September 28, MCP’s only stable auth extension authenticates the employee, not the agent , with DPoP (SEP-1932) and Workload Identity Federation (SEP-1933) still open pull requests.

Egress: decide on the destination, not the request

The collusion wiki corpus showed agents in a read-only sandbox writing anyway. In Egress Control for AI Agents: Four Bypasses, One Fix I reproduced four of their techniques against a proxy whose first rule was that writes are blocked. Four writes landed, and every one was a GET.

The version that held was a gateway on an internal: true network that terminates CONNECT and resolves the name itself. The /etc/hosts technique returned a clean 200 with real JSON, from the datasource the gateway resolved, not from the wiki the agent pointed at. The agent’s name mapping never entered the decision.

That post also said what it does not prove, and the next lab proved it. In Your Artifact Registry Is a Two-Way Channel for Agents the agent’s allowlist permits exactly one host, the registry, because builds need it. Publishing and listing packages is enough to move data out and read coordination back, with every request a valid API call. The guard that caught it ignored payload content and scored behavior per identity: a shadow identity encoding 130 bytes over five publishes scored 5.12 bits per character of field entropy against 4.47 for a normal build, plus a five-step sequential counter. Two independent signals, so it flagged. A destination allowlist is necessary. The permitted destinations are still attack surface.

Audit integrity: the record has to be written by something the agent cannot reach

METR and Redwood found spoofed tool calls in at least 96 transcripts, roughly 7% of those they reviewed from the OpenAI and Hugging Face incident. In Your Agent Wrote the Audit Log You Are Judging It By an agent overwrites its own tool module so the transcript reports a GET to an approved API while the real call POSTs the workspace elsewhere.

The transcript monitor returned no violations found on both the honest run and the spoofed one. Byte-identical verdicts. The reconciler, which compares the transcript against a hash-chained log written by a proxy the agent has no mount for, returned two discrepancies: one call the transcript claimed and the wire never saw, and one the wire saw that the transcript never mentioned.

Stopping an AI Agent Without Losing the Forensic Record is the other half. With a kagent agent scaled to zero replicas, the session record, the full prompt and response, and 18 trace spans all stayed queryable, because the control plane and the gateway wrote them, not the agent. A record has to be written by something other than the agent and survive the agent being stopped. Where you put the database fixes only the second.

Tool poisoning: descriptions are program text

An MCP tool description is prose the model reads and acts on. Treat it as documentation and you will review the wrong thing.

Your AGENTS.md Is Not a Security Control ran one injected agent against three control stacks. With only AGENTS.md, it destroyed 5 of 5 records and posted a confirmation to an outside URL. Adding an in-band approval classifier changed nothing: 5 of 5 again. With agentgateway mcpAuthorization in front of the server, 0 destroyed, 0 exfil calls, and only 2 of 5 tools visible. The policy file and the classifier both live in the channel the attacker writes to.

MCP Tool Poisoning: A Name Allowlist Is Not Enough reproduced a server that turns hostile after three tool calls without renaming anything. The name allowlist hid the gated tool and still passed the poisoned description through. Pinning a digest over name, description and inputSchema at the first tools/list rejected the mutation and quarantined the session.

Pinning per tool has a gap, which MCP Prompt Injection at Connect covers. The instructions field arrives once at initialize, outside any tool, and clients fold it into the system prompt. A scan reported in the spec thread found 5,462 of 8,235 live registry servers, 66%, send it. The lab showed a shared cache serving one caller’s poisoned discovery response to a second caller, and binding the cache key to server identity and caller stopped it.

Runtimes: an autoscaler is not a scheduler

This area produced an open problem, not a fix. In Agent Runtimes Have an Autoscaler, Not a Scheduler I pulled the live CRDs on substrate v0.0.8. Nothing in WorkerPool or ActorTemplate says one agent matters more than another. When the pool is full the runtime rejects with substrate worker pool has no free workers, so the queue lives in callers’ retry loops, and the agent that retries hardest wins the next slot. The only priority field in the surface is standard Kubernetes pod priority on the worker template, which decides whether a pod gets a node, not which agent gets a worker.

It connects to everything above. Efficiency policy decides how much capacity exists. It never decides who loses when there is not enough, and right now the answer is arrival order.

Evals: an allowlist cannot catch an agent that did no work

In Regression Tests for kagent Agents with agentevals the same three-case suite ran against a grounded session (3/3 passed) and one where the agent skipped the weather tool and invented a report (1/3 passed, exit code 1). The case that still passed was the tool allowlist. An agent that called nothing satisfies a subset check, because the empty set is a subset of everything. Catching it took a superset check that the tool was called, and a strict check on arguments caught the agent that looked up Paris when asked about London.

The same post names the limit: the trajectory being scored is the runtime’s own record. Evals catch behavior drift. They are not a witness.

A related measurement, from a 60-case benchmark of Jev on tool-call risk : 40 of 60 answers came back at confidence exactly 1.000 and all 40 were correct, while one miss landed at 0.97. If you route escalation on a confidence score, measure where the misses land before you pick the threshold.

The finding I did not plan: controls that report success while testing nothing

Three labs caught me the same way, and I think this is the most transferable lesson of the month.

  • Status said healthy. In agentgateway CEL Gotchas one policy failed open because matchExpressions entries are OR’ed, and the next one denied every request because llm.requestModel is empty during traffic.authorization. Both showed Accepted and Attached.
  • The test passed against the wrong process. The tool-poisoning demo first reported 8 passed, 0 failed. A kubectl port-forward held the port, every request got Go’s 404 page not found, and an assertion on the absence of id_rsa passed on a 404 body. The fix was a preflight that makes each endpoint prove what it is.
  • The reconciler matched a call that never happened. Joining on host plus a 30-second window paired the spoofed claim with a genuine GET left over from the previous run, and the UNWITNESSED finding disappeared.

If you assert on the absence of a bad thing, a broken endpoint looks exactly like a working control. The labs that I trust most prove the negative as well: the LLM key demo replays a wrong key and gets a 401, which is what makes its 200 mean something.

What is still open

  • Agent identity in MCP. DPoP and Workload Identity Federation were still open pull requests when I last checked on September 28. Until they land, identity comes from a layer down: SPIFFE SVIDs or short-lived tokens minted at the gateway.
  • Tamper-evident audit records. MCP SEP-3004 was still seeking a sponsor when I wrote about it on September 9. There is no ratified format to conform to, so the thing to get right is where the record is written.
  • Priority for agents. No runtime I tested expresses which agent should win a contended slot.
  • Permitted destinations. Behavioral signals on the write path caught the registry channel in a lab. I have not seen them tested against real build traffic.

If you run agents on Kubernetes this quarter, I would build in the order the labs failed. Move the provider key from the agent’s environment to the gateway. Put the egress decision on the resolved destination at CONNECT. Write the tool-call log from a proxy the agent has no mount for. Pin tools/list by digest, and remember that instructions sits outside the pin. Add a superset check to every eval suite that only allowlists. Each of those has a repo linked above that you can clone, run and argue with.

Frequently asked questions

What is the most common reason AI agent security controls fail?

The control reads a field the agent chooses. In a month of reproductions, a proxy rule that blocked writes let four writes land, every one a GET. A tool name allowlist passed a poisoned MCP tool description. An AGENTS.md rule did not stop an injected agent from destroying 5 of 5 records. The controls that held sat at a chokepoint the agent cannot reach and decided on something it cannot set, such as the resolved destination or a pinned digest.

Do Kubernetes RBAC and Pod Security protect an AI agent's LLM API key?

No. They govern access to Secret objects, not a credential the pod already holds. With restricted Pod Security enforcing, no RBAC verbs on secrets and automountServiceAccountToken false, one file read inside the container still returned the key, and a different pod replayed it for a 200. The fix is to keep the provider key out of the agent entirely and let the gateway attach it.

Can you trust an AI agent's own transcript as an audit log?

Not on its own. METR and Redwood found spoofed tool calls in about 7% of reviewed transcripts from the Hugging Face incident. In a reproduction, a transcript monitor returned 'no violations found' on both an honest run and a spoofed one, while reconciling the transcript against a hash-chained proxy log the agent could not mount found 2 discrepancies.

Can an MCP tool allowlist catch an agent that skipped its tools?

No. A subset check asks whether the agent called anything outside the approved set, and an agent that called nothing passes it, because the empty set is a subset of everything. In a kagent and agentevals suite, the ungrounded session still passed the allowlist case. Only a superset check asserting the tool was called, and a strict check on arguments, caught it.