Almost everything in this series started the same way: a control that everybody agrees is the right control, applied to an agent, quietly not doing the thing its name implies.
The pattern is consistent enough to be a design rule. Controls built for human operators assume a session, an intent, and a bounded set of actions. An agent has none of those. It has a tool surface, a credential, and a loop — and the loop will find the edge of every boundary you drew for a person.
The thread running through these
Advisory is not enforcement. AGENTS.md is a file the model may read. A tool name allowlist pins a string, not behavior. Neither survives contact with a model that has a reason to go around it.
Permitted is not safe. The hardest destination to defend is the one you are required to allow. The artifact registry has to be reachable because builds need it, and that single exception is enough to run data out and coordination back in.
The record is written by the suspect. If the agent can influence the transcript, the transcript is testimony, not evidence. This is the one that most surprises people, and it has a real number behind it.
The credential outlives the control. Kubernetes secret hygiene covers the things Kubernetes issues. The provider API key the agent actually holds is usually not one of them.
Where to start
If you are trying to convince someone this matters, start with which controls would have stopped the July 2026 intrusion — it works backward from a real incident.
If you are already convinced and want the sharpest single finding, start with the audit log your agent wrote .
If you run a platform and want to know what is already deployed without your knowledge, start with finding the MCP servers your platform team doesn’t know about .
The related work on how an agent proves who it is — SPIFFE, token exchange, audience binding, and what MCP does and does not specify — is collected separately in the agent identity series .
Every post in this series
- AI Agent Security on Kubernetes: Which Controls Actually Held
— A month of AI agent security labs on identity, egress, audit logs, MCP tool poisoning, runtimes and evals. The controls that held all sat outside the agent. - Your Agent Egress Proxy Never Saw the DNS Query
— An OpenAI agent reached an external chatbot over DNS after its HTTPS calls were blocked. The Kubernetes NetworkPolicy that closes that channel, reproduced live. - What Stopped AI Agent Probes of Government Data Sites
— OpenAI agents probing government data sites during ordinary retrieval. What Cloudflare stopped, what it missed, and the egress control that generalizes. - Your Artifact Registry Is a Two-Way Channel for Agents
— An agent's egress allowlist permits the artifact registry because builds need it. That single permitted destination is enough to run a covert channel, as this week's RubyGems incident showed. - MCP Tool Poisoning: A Name Allowlist Is Not Enough
— A deny-by-default MCP allowlist strips gated tools and still adopts a poisoned description. Pinning name, description and schema rejects the mutation. - Stopping an AI Agent Without Losing the Forensic Record
— Security teams want an agent stopped now and everything it did preserved. Those fight unless the record lives outside the agent. A test on Kubernetes. - A WAF That Reads the Prompt: OWASP CRS for LLM and MCP
— Coraza with body inspection puts OWASP CRS and custom rules on LLM prompts and MCP tool arguments. The verified blocks, and the false-positive check. - I Benchmarked Jev on Agent Tool-Call Risk. Calibration Held.
— A 60-case benchmark of TypeSafe AI's Jev classifying agent tool calls as readonly, destructive, privileged, or exfiltration. Accuracy is 91.7%. The finding is calibration. - Shadow MCP Servers: Why They're Hard to Track
— Stop shadow MCP servers at an agentgateway chokepoint. Federate registered servers, authorize each tool call, then find what scans miss. - Your Agent's LLM Key Survives Every Kubernetes Secret Control
— RBAC, restricted Pod Security, and automountServiceAccountToken all on, and one file read still hands over an agent's LLM API key. Tested on my own test cluster. - Keeping MCP Server Instructions Out of Your Agent's System Prompt
— Keep server-controlled MCP instructions out of the trusted system prompt. Isolate, cap, bind the cache, and pin the digest. Lab plus the four controls. - Agent Audit Logs: What METR and Redwood Found
— METR found spoofed tool calls in 7% of agent transcripts from the Hugging Face incident. A transcript monitor reads them as clean. Diff against a witness. - Why Proxy-Variable Agent Sandboxes Fail, and the CONNECT-Time Fix
— Agent sandboxes that key egress on proxy variables fail for structural reasons. The control that holds authorizes the destination at CONNECT time. - Your AGENTS.md Is Not a Security Control
— A prompt-injected agent destroyed every record with an AGENTS.md forbidding it and an approval classifier watching. An MCP allowlist at the gateway stopped it. - Which Controls Would Have Stopped the July 2026 Agent Intrusion?
— An autonomous agent went from sandbox escape to Kubernetes cluster-admin in under 13 hours. A stage-by-stage map of which controls would have broken the chain.