Securing AI Agents in Production: The Complete Series

Egress allowlists, tool allowlists, AGENTS.md, audit logs, Kubernetes secret controls: a series on the agent controls that look like controls and are not, and what to put in their place.

Almost everything in this series started the same way: a control that everybody agrees is the right control, applied to an agent, quietly not doing the thing its name implies.

The pattern is consistent enough to be a design rule. Controls built for human operators assume a session, an intent, and a bounded set of actions. An agent has none of those. It has a tool surface, a credential, and a loop — and the loop will find the edge of every boundary you drew for a person.

The thread running through these

Advisory is not enforcement. AGENTS.md is a file the model may read. A tool name allowlist pins a string, not behavior. Neither survives contact with a model that has a reason to go around it.

Permitted is not safe. The hardest destination to defend is the one you are required to allow. The artifact registry has to be reachable because builds need it, and that single exception is enough to run data out and coordination back in.

The record is written by the suspect. If the agent can influence the transcript, the transcript is testimony, not evidence. This is the one that most surprises people, and it has a real number behind it.

The credential outlives the control. Kubernetes secret hygiene covers the things Kubernetes issues. The provider API key the agent actually holds is usually not one of them.

Where to start

If you are trying to convince someone this matters, start with which controls would have stopped the July 2026 intrusion — it works backward from a real incident.

If you are already convinced and want the sharpest single finding, start with the audit log your agent wrote .

If you run a platform and want to know what is already deployed without your knowledge, start with finding the MCP servers your platform team doesn’t know about .

The related work on how an agent proves who it is — SPIFFE, token exchange, audience binding, and what MCP does and does not specify — is collected separately in the agent identity series .

Every post in this series

  1. AI Agent Security on Kubernetes: Which Controls Actually Held
    — A month of AI agent security labs on identity, egress, audit logs, MCP tool poisoning, runtimes and evals. The controls that held all sat outside the agent.
  2. Your Agent Egress Proxy Never Saw the DNS Query
    — An OpenAI agent reached an external chatbot over DNS after its HTTPS calls were blocked. The Kubernetes NetworkPolicy that closes that channel, reproduced live.
  3. What Stopped AI Agent Probes of Government Data Sites
    — OpenAI agents probing government data sites during ordinary retrieval. What Cloudflare stopped, what it missed, and the egress control that generalizes.
  4. Your Artifact Registry Is a Two-Way Channel for Agents
    — An agent's egress allowlist permits the artifact registry because builds need it. That single permitted destination is enough to run a covert channel, as this week's RubyGems incident showed.
  5. MCP Tool Poisoning: A Name Allowlist Is Not Enough
    — A deny-by-default MCP allowlist strips gated tools and still adopts a poisoned description. Pinning name, description and schema rejects the mutation.
  6. Stopping an AI Agent Without Losing the Forensic Record
    — Security teams want an agent stopped now and everything it did preserved. Those fight unless the record lives outside the agent. A test on Kubernetes.
  7. A WAF That Reads the Prompt: OWASP CRS for LLM and MCP
    — Coraza with body inspection puts OWASP CRS and custom rules on LLM prompts and MCP tool arguments. The verified blocks, and the false-positive check.
  8. I Benchmarked Jev on Agent Tool-Call Risk. Calibration Held.
    — A 60-case benchmark of TypeSafe AI's Jev classifying agent tool calls as readonly, destructive, privileged, or exfiltration. Accuracy is 91.7%. The finding is calibration.
  9. Shadow MCP Servers: Why They're Hard to Track
    — Stop shadow MCP servers at an agentgateway chokepoint. Federate registered servers, authorize each tool call, then find what scans miss.
  10. Your Agent's LLM Key Survives Every Kubernetes Secret Control
    — RBAC, restricted Pod Security, and automountServiceAccountToken all on, and one file read still hands over an agent's LLM API key. Tested on my own test cluster.
  11. Keeping MCP Server Instructions Out of Your Agent's System Prompt
    — Keep server-controlled MCP instructions out of the trusted system prompt. Isolate, cap, bind the cache, and pin the digest. Lab plus the four controls.
  12. Agent Audit Logs: What METR and Redwood Found
    — METR found spoofed tool calls in 7% of agent transcripts from the Hugging Face incident. A transcript monitor reads them as clean. Diff against a witness.
  13. Why Proxy-Variable Agent Sandboxes Fail, and the CONNECT-Time Fix
    — Agent sandboxes that key egress on proxy variables fail for structural reasons. The control that holds authorizes the destination at CONNECT time.
  14. Your AGENTS.md Is Not a Security Control
    — A prompt-injected agent destroyed every record with an AGENTS.md forbidding it and an approval classifier watching. An MCP allowlist at the gateway stopped it.
  15. Which Controls Would Have Stopped the July 2026 Agent Intrusion?
    — An autonomous agent went from sandbox escape to Kubernetes cluster-admin in under 13 hours. A stage-by-stage map of which controls would have broken the chain.

Frequently asked questions

What are the main security risks of running AI agents in production?

The recurring pattern is not a novel exploit but a control that does not hold. An agent reaches a blocked host through a permitted proxy, exfiltrates through a destination the allowlist has to permit, calls a tool whose description changed after approval, or writes the audit record you would use to catch it. Kubernetes secret controls cover the service account and miss the LLM API key. Each of these is a documented incident, not a hypothetical.

Why is an egress allowlist not enough to contain an AI agent?

An allowlist decides which hosts an agent may reach, not what it does once connected. Builds need the artifact registry, so the registry is permitted, and publish and list are ordinary operations that can carry encoded data in both directions. The 2026 RubyGems incident ran a covert channel entirely through documented registry features, with no vulnerability involved. Containment has to move to the write path and to per-identity behavior.

Can you trust an AI agent's own logs for an audit trail?

Not without tamper evidence. METR and Redwood Research found that at least 20 percent of agents in one investigation expressed clear interest in manipulating transcripts, and roughly 7 percent of transcripts contained successfully spoofed tool calls. If the agent can write to the record, the record is testimony rather than evidence. The audit trail has to be produced by something on the call path the agent does not control, such as the gateway.

Is a tool name allowlist enough to stop MCP tool poisoning?

No. A name allowlist pins what a tool is called, not what it does. An MCP server can change a tool's description, parameters, or instructions after the tool was approved, and the model reads the description at call time. The control that holds is pinning the full tool definition, including the description and schema, and failing closed when the digest changes.