Three replicas of an MCP server, fronted by agentgateway , with no session affinity configured anywhere. A four-call shopping cart session (create cart, add two items, check out) completes correctly across different pods, including mid-session through a full rolling restart of the Deployment. That is the practical Kubernetes payoff of the MCP 2026-07-28 spec revision , and stateless-mcp-scale is the repo that proves it instead of just describing it.
Simon Willison wrote about this spec revision the week it landed, in Stateless MCP has recaptured my interest
, which picked up 386 points on Hacker News
. His post explains the wire-format change clearly: the old two-request dance (initialize to get an Mcp-Session-Id, then a second request using it) collapses into one self-contained request. What his post doesn’t cover, because it isn’t about Kubernetes operations, is what that change actually buys an infrastructure team: the end of a specific, annoying scaling constraint. This post builds the smallest cluster that shows it.
What the old spec required
Before 2026-07-28, an MCP client had to call initialize, get back an Mcp-Session-Id header, and send that header on every subsequent request in the session. The server was expected to keep state tied to that session ID in memory. On Kubernetes, that meant whichever pod handled initialize had to handle every later request in the session too, which meant either sessionAffinity: ClientIP on the Service or cookie-based affinity in front of it. Session affinity isn’t exotic, but it has a cost: it fights horizontal pod autoscaling (a new replica only gets new sessions, never a share of existing ones), and it turns a rolling restart or a scale-down into a reliability event for whatever session was pinned to the terminated pod.
What the new spec removes
The spec’s own announcement
states the change plainly: “Each request now travels on its own, carrying its protocol version, client identity, and client capabilities in _meta,” which means “any request [can land] on any server instance behind a plain round-robin load balancer without needing shared storage.” The initialize/initialized handshake and the Mcp-Session-Id header are both retired for this mode, not made optional. Where a server genuinely needs state to outlive a single call, the spec’s guidance is direct: “mint an explicit handle from a tool and have the model pass it back as an argument.” State moves from the server’s memory into the request itself.
agentgateway’s own v1.4.0 release notes
describe the same mechanism from the gateway’s side, under “MCP protocol 2026-07-28 support”: support for “stateful and stateless servers, including closing the SEP-2575 server-stateless conformance gap and skipping the synthetic initialize handshake for modern requests.” That statefulMode setting is what this demo turns on.
What I built
stateless-mcp-scale has three pieces:
- A minimal MCP server (
server/stateless_mcp_server.py, Python stdlib only, no framework) implementinginitialize,tools/list, andtools/callover Streamable HTTP. It exposes three tools,cart_create,cart_add_item, andcart_checkout, whose state lives entirely in an opaquecartTokenthe caller holds, not in the process. Every response also reportsservedBy.podandservedBy.podIPso the demo can show which replica answered. - A 3-replica Kubernetes Deployment of that server, behind a plain
ClusterIPService. Nothing setssessionAffinity, so it stays at the Kubernetes default,None. - agentgateway v1.5.0 in front of it, configured with
statefulMode: stateless:
gateways:
default:
port: 3000
mcp:
gateways:
- default
statefulMode: stateless
targets:
- name: stateless-mcp-demo
mcp:
host: stateless-mcp-demo.stateless-mcp-scale.svc.cluster.local
port: 8080
path: /
statefulMode: stateless is documented in agentgateway’s config schema
as the choice between keeping “a persistent session across requests (Stateful)” or creating “one per request (Stateless).” I validated this exact file before deploying it:
docker run --rm -v "$PWD/agentgateway:/config" \
cr.agentgateway.dev/agentgateway:v1.5.0 \
--validate-only -f /config/agentgateway.yaml
Configuration is valid!
agentgateway itself is also exposed through a plain NodePort Service with no sessionAffinity set. Under the old spec, this entire setup, a client bouncing between replicas with no pinning at either hop, would have broken the first time a session outlived the connection that started it. Under the new one, it’s the default.
The wire format, the hard way
None of this is documented end to end yet; agentgateway’s own release notes admit as much (“most of it is not yet covered by a dedicated guide”). I found the exact requirements by sending bare requests and reading agentgateway’s error messages, which turned out to be precise enough to iterate against directly:
400 mcp: invalid MCP protocol version header
Fixed by adding MCP-Protocol-Version: 2026-07-28 as an HTTP header, not just inside the JSON-RPC body.
400 invalid request parameters: _meta.protocolVersion is required for modern requests
Fixed by adding _meta to the JSON-RPC params, with the key spelled the reverse-DNS-qualified way, io.modelcontextprotocol/protocolVersion, not a bare protocolVersion.
400 invalid MCP routing header: Mcp-Method
Fixed by adding Mcp-Method: tools/call (or tools/list, matching the JSON-RPC method) as its own HTTP header, plus Mcp-Name: cart_create for tool calls. These are agentgateway’s own routing headers, separate from the JSON-RPC envelope.
And trying to send a standard initialize call first gets:
{"jsonrpc":"2.0","id":1,"error":{"code":-32601,"message":"method not found: initialize"}}
With statefulMode: stateless, agentgateway rejects the handshake outright. That matches the spec’s own framing: a modern, stateless client doesn’t send one. At no point in any of this does agentgateway send or expect an Mcp-Session-Id, to the client or to the backend. scripts/mcp_client.py in the repo encodes all four requirements in one place so nobody else has to rediscover them by reading error strings.
The proof
make up builds the demo image, creates a kind cluster, loads both images into it, and applies the namespace, the 3-replica Deployment, and agentgateway, waiting on each rollout. make demo runs scripts/demo.py, which does three things against the live cluster and asserts on the results. This is the unedited output from an actual clean run:
=== 1. Fan-out: 9 independent tools/call requests, no session affinity configured ===
request 1: served by stateless-mcp-demo-757875b499-z2zsg
request 2: served by stateless-mcp-demo-757875b499-w6mpr
request 3: served by stateless-mcp-demo-757875b499-z2zsg
request 4: served by stateless-mcp-demo-757875b499-w6mpr
request 5: served by stateless-mcp-demo-757875b499-z2zsg
request 6: served by stateless-mcp-demo-757875b499-w6mpr
request 7: served by stateless-mcp-demo-757875b499-z2zsg
request 8: served by stateless-mcp-demo-757875b499-z2zsg
request 9: served by stateless-mcp-demo-757875b499-z2zsg
-> 2 distinct replica(s) answered 9 requests: ['stateless-mcp-demo-757875b499-w6mpr', 'stateless-mcp-demo-757875b499-z2zsg']
=== 2. One logical session, four separate requests, zero shared server memory ===
cart_create -> pod=stateless-mcp-demo-757875b499-z2zsg cartToken(len=108)
cart_add_item -> pod=stateless-mcp-demo-757875b499-w6mpr items=['keyboard']
cart_add_item -> pod=stateless-mcp-demo-757875b499-w6mpr items=['keyboard', 'mouse']
cart_checkout -> pod=stateless-mcp-demo-757875b499-w6mpr total=$65.00 status=confirmed
-> session correct across every hop, no Mcp-Session-Id, no sticky routing.
=== 3. Bonus: rolling restart of every replica, mid-session ===
cart_create (before restart) -> pod=stateless-mcp-demo-757875b499-z2zsg
triggered: kubectl rollout restart deployment/stateless-mcp-demo
cart_add_item during rollout (1/5) -> pod=stateless-mcp-demo-757875b499-n9s5r
cart_add_item during rollout (2/5) -> pod=stateless-mcp-demo-757875b499-z2zsg
cart_add_item during rollout (3/5) -> pod=stateless-mcp-demo-7674b87667-9vvrw
cart_add_item during rollout (4/5) -> pod=stateless-mcp-demo-7674b87667-9vvrw
cart_add_item during rollout (5/5) -> pod=stateless-mcp-demo-7674b87667-rp9fm
cart_checkout (after restart) -> pod=stateless-mcp-demo-7674b87667-9vvrw total=$5.00 status=confirmed
-> cart survived a full rolling restart: created on stateless-mcp-demo-757875b499-z2zsg, checked out on stateless-mcp-demo-7674b87667-9vvrw (that pod didn't exist when the cart was created).
All assertions passed.
Step 1 landed on 2 of the 3 replicas in this particular 9-request sample; a separate run during development hit all 3 of 3 in the first 6. The exact ratio isn’t the point. The point is that no two consecutive calls were pinned together by any mechanism, and every call still succeeded. Step 3 is the one I’d reread twice: cart_checkout closes out a cart that was created on a pod from the Deployment’s previous ReplicaSet, a pod that no longer exists by the time checkout runs. The replica that answers checkout has never seen that cart in its own memory, and decodes it correctly anyway, because the cart’s state was never in that memory. Under the pre-2026-07-28 model, this exact sequence, no affinity plus a rolling restart mid-session, is close to the textbook way to break an MCP client.
Correcting the premise
The angle I started with was that agentgateway shipped support for the new spec “the same day it landed.” Checking gh release list -R agentgateway/agentgateway and gh release view v1.4.0 shows that’s not quite right, and the real timeline is more interesting. agentgateway v1.4.0, which added full 2026-07-28 support, published 2026-07-27 at 17:51 UTC, one day before the spec’s dated revision went final. It was built against the spec’s release candidate
, public since May 21, 2026, giving implementers over two months of lead time. v1.4.1 followed on 2026-07-29 with a round of MCP compatibility fixes, including validating upstream MCP responses against their expected type and no longer logging synthetic session identifiers as if they were real Mcp-Session-Id values. Willison’s post and the Hacker News discussion both came a few days after that, on July 31 and August 1. “Day one” support is still accurate in spirit; “same day” was the wrong detail, and the actual sequence (built against the RC, shipped a day early, patched two days later) is a better story about how a spec revision this size actually gets implemented.
I’ve written before about federating MCP servers behind agentgateway for multi-tenant deployments , where the axis of scale is different customers getting different tools from the same gateway. This post is a different axis: identical replicas of the same server, and the question is whether any of them can answer any request. Under the new spec, the answer is yes, with nothing extra configured to make it true.
The full repo, including the server, the Kubernetes manifests, the agentgateway config, and the scripts that produced the output above, is at github.com/themsquared/stateless-mcp-scale
. make up && make demo reproduces all three experiments on a kind cluster. The next thing worth testing is a multi-node cluster with a real HPA scaling the MCP server under load, to see the same no-affinity property hold when pods are moving across machines, not just across ReplicaSets on one node.