Stateless MCP on Kubernetes: No More Sticky Sessions

A 3-replica MCP server behind agentgateway, no session affinity, survives a rolling restart mid-session. What the MCP 2026-07-28 spec changed, and the proof.

Three replicas of an MCP server, fronted by agentgateway , with no session affinity configured anywhere. A four-call shopping cart session (create cart, add two items, check out) completes correctly across different pods, including mid-session through a full rolling restart of the Deployment. That is the practical Kubernetes payoff of the MCP 2026-07-28 spec revision , and stateless-mcp-scale is the repo that proves it instead of just describing it.

Simon Willison wrote about this spec revision the week it landed, in Stateless MCP has recaptured my interest , which picked up 386 points on Hacker News . His post explains the wire-format change clearly: the old two-request dance (initialize to get an Mcp-Session-Id, then a second request using it) collapses into one self-contained request. What his post doesn’t cover, because it isn’t about Kubernetes operations, is what that change actually buys an infrastructure team: the end of a specific, annoying scaling constraint. This post builds the smallest cluster that shows it.

What the old spec required

Before 2026-07-28, an MCP client had to call initialize, get back an Mcp-Session-Id header, and send that header on every subsequent request in the session. The server was expected to keep state tied to that session ID in memory. On Kubernetes, that meant whichever pod handled initialize had to handle every later request in the session too, which meant either sessionAffinity: ClientIP on the Service or cookie-based affinity in front of it. Session affinity isn’t exotic, but it has a cost: it fights horizontal pod autoscaling (a new replica only gets new sessions, never a share of existing ones), and it turns a rolling restart or a scale-down into a reliability event for whatever session was pinned to the terminated pod.

What the new spec removes

The spec’s own announcement states the change plainly: “Each request now travels on its own, carrying its protocol version, client identity, and client capabilities in _meta,” which means “any request [can land] on any server instance behind a plain round-robin load balancer without needing shared storage.” The initialize/initialized handshake and the Mcp-Session-Id header are both retired for this mode, not made optional. Where a server genuinely needs state to outlive a single call, the spec’s guidance is direct: “mint an explicit handle from a tool and have the model pass it back as an argument.” State moves from the server’s memory into the request itself.

agentgateway’s own v1.4.0 release notes describe the same mechanism from the gateway’s side, under “MCP protocol 2026-07-28 support”: support for “stateful and stateless servers, including closing the SEP-2575 server-stateless conformance gap and skipping the synthetic initialize handshake for modern requests.” That statefulMode setting is what this demo turns on.

What I built

stateless-mcp-scale has three pieces:

  1. A minimal MCP server (server/stateless_mcp_server.py, Python stdlib only, no framework) implementing initialize, tools/list, and tools/call over Streamable HTTP. It exposes three tools, cart_create, cart_add_item, and cart_checkout, whose state lives entirely in an opaque cartToken the caller holds, not in the process. Every response also reports servedBy.pod and servedBy.podIP so the demo can show which replica answered.
  2. A 3-replica Kubernetes Deployment of that server, behind a plain ClusterIP Service. Nothing sets sessionAffinity, so it stays at the Kubernetes default, None.
  3. agentgateway v1.5.0 in front of it, configured with statefulMode: stateless:
gateways:
  default:
    port: 3000

mcp:
  gateways:
    - default
  statefulMode: stateless
  targets:
    - name: stateless-mcp-demo
      mcp:
        host: stateless-mcp-demo.stateless-mcp-scale.svc.cluster.local
        port: 8080
        path: /

statefulMode: stateless is documented in agentgateway’s config schema as the choice between keeping “a persistent session across requests (Stateful)” or creating “one per request (Stateless).” I validated this exact file before deploying it:

docker run --rm -v "$PWD/agentgateway:/config" \
  cr.agentgateway.dev/agentgateway:v1.5.0 \
  --validate-only -f /config/agentgateway.yaml
Configuration is valid!

agentgateway itself is also exposed through a plain NodePort Service with no sessionAffinity set. Under the old spec, this entire setup, a client bouncing between replicas with no pinning at either hop, would have broken the first time a session outlived the connection that started it. Under the new one, it’s the default.

The wire format, the hard way

None of this is documented end to end yet; agentgateway’s own release notes admit as much (“most of it is not yet covered by a dedicated guide”). I found the exact requirements by sending bare requests and reading agentgateway’s error messages, which turned out to be precise enough to iterate against directly:

400 mcp: invalid MCP protocol version header

Fixed by adding MCP-Protocol-Version: 2026-07-28 as an HTTP header, not just inside the JSON-RPC body.

400 invalid request parameters: _meta.protocolVersion is required for modern requests

Fixed by adding _meta to the JSON-RPC params, with the key spelled the reverse-DNS-qualified way, io.modelcontextprotocol/protocolVersion, not a bare protocolVersion.

400 invalid MCP routing header: Mcp-Method

Fixed by adding Mcp-Method: tools/call (or tools/list, matching the JSON-RPC method) as its own HTTP header, plus Mcp-Name: cart_create for tool calls. These are agentgateway’s own routing headers, separate from the JSON-RPC envelope.

And trying to send a standard initialize call first gets:

{"jsonrpc":"2.0","id":1,"error":{"code":-32601,"message":"method not found: initialize"}}

With statefulMode: stateless, agentgateway rejects the handshake outright. That matches the spec’s own framing: a modern, stateless client doesn’t send one. At no point in any of this does agentgateway send or expect an Mcp-Session-Id, to the client or to the backend. scripts/mcp_client.py in the repo encodes all four requirements in one place so nobody else has to rediscover them by reading error strings.

The proof

make up builds the demo image, creates a kind cluster, loads both images into it, and applies the namespace, the 3-replica Deployment, and agentgateway, waiting on each rollout. make demo runs scripts/demo.py, which does three things against the live cluster and asserts on the results. This is the unedited output from an actual clean run:

=== 1. Fan-out: 9 independent tools/call requests, no session affinity configured ===
  request 1: served by stateless-mcp-demo-757875b499-z2zsg
  request 2: served by stateless-mcp-demo-757875b499-w6mpr
  request 3: served by stateless-mcp-demo-757875b499-z2zsg
  request 4: served by stateless-mcp-demo-757875b499-w6mpr
  request 5: served by stateless-mcp-demo-757875b499-z2zsg
  request 6: served by stateless-mcp-demo-757875b499-w6mpr
  request 7: served by stateless-mcp-demo-757875b499-z2zsg
  request 8: served by stateless-mcp-demo-757875b499-z2zsg
  request 9: served by stateless-mcp-demo-757875b499-z2zsg
  -> 2 distinct replica(s) answered 9 requests: ['stateless-mcp-demo-757875b499-w6mpr', 'stateless-mcp-demo-757875b499-z2zsg']

=== 2. One logical session, four separate requests, zero shared server memory ===
  cart_create      -> pod=stateless-mcp-demo-757875b499-z2zsg      cartToken(len=108)
  cart_add_item    -> pod=stateless-mcp-demo-757875b499-w6mpr      items=['keyboard']
  cart_add_item    -> pod=stateless-mcp-demo-757875b499-w6mpr      items=['keyboard', 'mouse']
  cart_checkout    -> pod=stateless-mcp-demo-757875b499-w6mpr      total=$65.00 status=confirmed
  -> session correct across every hop, no Mcp-Session-Id, no sticky routing.

=== 3. Bonus: rolling restart of every replica, mid-session ===
  cart_create (before restart) -> pod=stateless-mcp-demo-757875b499-z2zsg
  triggered: kubectl rollout restart deployment/stateless-mcp-demo
  cart_add_item during rollout (1/5) -> pod=stateless-mcp-demo-757875b499-n9s5r
  cart_add_item during rollout (2/5) -> pod=stateless-mcp-demo-757875b499-z2zsg
  cart_add_item during rollout (3/5) -> pod=stateless-mcp-demo-7674b87667-9vvrw
  cart_add_item during rollout (4/5) -> pod=stateless-mcp-demo-7674b87667-9vvrw
  cart_add_item during rollout (5/5) -> pod=stateless-mcp-demo-7674b87667-rp9fm
  cart_checkout (after restart)  -> pod=stateless-mcp-demo-7674b87667-9vvrw total=$5.00 status=confirmed
  -> cart survived a full rolling restart: created on stateless-mcp-demo-757875b499-z2zsg, checked out on stateless-mcp-demo-7674b87667-9vvrw (that pod didn't exist when the cart was created).

All assertions passed.

Step 1 landed on 2 of the 3 replicas in this particular 9-request sample; a separate run during development hit all 3 of 3 in the first 6. The exact ratio isn’t the point. The point is that no two consecutive calls were pinned together by any mechanism, and every call still succeeded. Step 3 is the one I’d reread twice: cart_checkout closes out a cart that was created on a pod from the Deployment’s previous ReplicaSet, a pod that no longer exists by the time checkout runs. The replica that answers checkout has never seen that cart in its own memory, and decodes it correctly anyway, because the cart’s state was never in that memory. Under the pre-2026-07-28 model, this exact sequence, no affinity plus a rolling restart mid-session, is close to the textbook way to break an MCP client.

Correcting the premise

The angle I started with was that agentgateway shipped support for the new spec “the same day it landed.” Checking gh release list -R agentgateway/agentgateway and gh release view v1.4.0 shows that’s not quite right, and the real timeline is more interesting. agentgateway v1.4.0, which added full 2026-07-28 support, published 2026-07-27 at 17:51 UTC, one day before the spec’s dated revision went final. It was built against the spec’s release candidate , public since May 21, 2026, giving implementers over two months of lead time. v1.4.1 followed on 2026-07-29 with a round of MCP compatibility fixes, including validating upstream MCP responses against their expected type and no longer logging synthetic session identifiers as if they were real Mcp-Session-Id values. Willison’s post and the Hacker News discussion both came a few days after that, on July 31 and August 1. “Day one” support is still accurate in spirit; “same day” was the wrong detail, and the actual sequence (built against the RC, shipped a day early, patched two days later) is a better story about how a spec revision this size actually gets implemented.

I’ve written before about federating MCP servers behind agentgateway for multi-tenant deployments , where the axis of scale is different customers getting different tools from the same gateway. This post is a different axis: identical replicas of the same server, and the question is whether any of them can answer any request. Under the new spec, the answer is yes, with nothing extra configured to make it true.

The full repo, including the server, the Kubernetes manifests, the agentgateway config, and the scripts that produced the output above, is at github.com/themsquared/stateless-mcp-scale . make up && make demo reproduces all three experiments on a kind cluster. The next thing worth testing is a multi-node cluster with a real HPA scaling the MCP server under load, to see the same no-affinity property hold when pods are moving across machines, not just across ReplicaSets on one node.

Frequently asked questions

What changed in the MCP 2026-07-28 spec revision?

It retired the mandatory initialize/initialized handshake and the Mcp-Session-Id header that the earlier spec used to pin a client to one server process. Each request now carries its protocol version and client identity in its own _meta field, so any replica can answer any request. A server that needs state to survive between calls is expected to hand the caller an explicit token instead of keeping it in memory.

Does an MCP server need sticky sessions (session affinity) on Kubernetes?

Not if it speaks the 2026-07-28 stateless wire format. This post's demo runs a 3-replica MCP server behind agentgateway with sessionAffinity left at the Kubernetes default (None) on both Services, and a 4-call session completes correctly across different pods, including through a full rolling restart of the Deployment.

What HTTP headers does agentgateway require for a stateless MCP request?

Add MCP-Protocol-Version: 2026-07-28 as its own HTTP header, not just in the JSON-RPC body. Add Mcp-Method naming the JSON-RPC method, plus Mcp-Name for tools/call, as agentgateway's own routing headers, separate from the JSON-RPC envelope. The JSON-RPC params also need a _meta key named io.modelcontextprotocol/protocolVersion. Mcp-Session-Id is never sent or required.

Did agentgateway support the new MCP spec the day it shipped?

Almost: agentgateway v1.4.0, which added full 2026-07-28 support, published on 2026-07-27, one day before the spec's own dated revision went final. It was built against the spec's release candidate, public since May 21, 2026. v1.4.1 followed on 2026-07-29 with MCP compatibility bug fixes.