# Stateless MCP on Kubernetes: No More Sticky Sessions

> A 3-replica MCP server behind agentgateway, no session affinity, survives a rolling restart mid-session. What the MCP 2026-07-28 spec changed, and the proof.

- Canonical URL: https://webofmike.com/stateless-mcp-scale/
- Author: Mike Moore (https://webofmike.com/about/)
- Published: 2026-10-03
- Last modified: 2026-10-03
- Tags: MCP, Kubernetes, AI Gateways, Platform Engineering
- Cite as: Mike Moore, "Stateless MCP on Kubernetes: No More Sticky Sessions", Web of Mike (webofmike.com), 2026-10-03. https://webofmike.com/stateless-mcp-scale/


Three replicas of an MCP server, fronted by [agentgateway](https://agentgateway.dev), with no session affinity configured anywhere. A four-call shopping cart session (create cart, add two items, check out) completes correctly across different pods, including mid-session through a full rolling restart of the Deployment. That is the practical Kubernetes payoff of the MCP [2026-07-28 spec revision](https://blog.modelcontextprotocol.io/posts/2026-07-28/), and [stateless-mcp-scale](https://github.com/themsquared/stateless-mcp-scale) is the repo that proves it instead of just describing it.

Simon Willison wrote about this spec revision the week it landed, in [Stateless MCP has recaptured my interest](https://simonwillison.net/2026/Jul/31/stateless-mcp/), which picked up 386 points on [Hacker News](https://news.ycombinator.com/item?id=49131438). His post explains the wire-format change clearly: the old two-request dance (`initialize` to get an `Mcp-Session-Id`, then a second request using it) collapses into one self-contained request. What his post doesn't cover, because it isn't about Kubernetes operations, is what that change actually buys an infrastructure team: the end of a specific, annoying scaling constraint. This post builds the smallest cluster that shows it.

## What the old spec required

Before 2026-07-28, an MCP client had to call `initialize`, get back an `Mcp-Session-Id` header, and send that header on every subsequent request in the session. The server was expected to keep state tied to that session ID in memory. On Kubernetes, that meant whichever pod handled `initialize` had to handle every later request in the session too, which meant either `sessionAffinity: ClientIP` on the Service or cookie-based affinity in front of it. Session affinity isn't exotic, but it has a cost: it fights horizontal pod autoscaling (a new replica only gets new sessions, never a share of existing ones), and it turns a rolling restart or a scale-down into a reliability event for whatever session was pinned to the terminated pod.

## What the new spec removes

The [spec's own announcement](https://blog.modelcontextprotocol.io/posts/2026-07-28/) states the change plainly: "Each request now travels on its own, carrying its protocol version, client identity, and client capabilities in `_meta`," which means "any request [can land] on any server instance behind a plain round-robin load balancer without needing shared storage." The `initialize`/`initialized` handshake and the `Mcp-Session-Id` header are both retired for this mode, not made optional. Where a server genuinely needs state to outlive a single call, the spec's guidance is direct: "mint an explicit handle from a tool and have the model pass it back as an argument." State moves from the server's memory into the request itself.

agentgateway's own [v1.4.0 release notes](https://github.com/agentgateway/agentgateway/releases/tag/v1.4.0) describe the same mechanism from the gateway's side, under "MCP protocol 2026-07-28 support": support for "stateful and stateless servers, including closing the SEP-2575 server-stateless conformance gap and skipping the synthetic `initialize` handshake for modern requests." That `statefulMode` setting is what this demo turns on.

## What I built

[stateless-mcp-scale](https://github.com/themsquared/stateless-mcp-scale) has three pieces:

1. A minimal MCP server (`server/stateless_mcp_server.py`, Python stdlib only, no framework) implementing `initialize`, `tools/list`, and `tools/call` over Streamable HTTP. It exposes three tools, `cart_create`, `cart_add_item`, and `cart_checkout`, whose state lives entirely in an opaque `cartToken` the caller holds, not in the process. Every response also reports `servedBy.pod` and `servedBy.podIP` so the demo can show which replica answered.
2. A 3-replica Kubernetes Deployment of that server, behind a plain `ClusterIP` Service. Nothing sets `sessionAffinity`, so it stays at the Kubernetes default, `None`.
3. agentgateway v1.5.0 in front of it, configured with `statefulMode: stateless`:

```yaml
gateways:
  default:
    port: 3000

mcp:
  gateways:
    - default
  statefulMode: stateless
  targets:
    - name: stateless-mcp-demo
      mcp:
        host: stateless-mcp-demo.stateless-mcp-scale.svc.cluster.local
        port: 8080
        path: /
```

`statefulMode: stateless` is documented in agentgateway's [config schema](https://raw.githubusercontent.com/agentgateway/agentgateway/main/schema/config.json) as the choice between keeping "a persistent session across requests (Stateful)" or creating "one per request (Stateless)." I validated this exact file before deploying it:

```bash
docker run --rm -v "$PWD/agentgateway:/config" \
  cr.agentgateway.dev/agentgateway:v1.5.0 \
  --validate-only -f /config/agentgateway.yaml
```

```
Configuration is valid!
```

agentgateway itself is also exposed through a plain `NodePort` Service with no `sessionAffinity` set. Under the old spec, this entire setup, a client bouncing between replicas with no pinning at either hop, would have broken the first time a session outlived the connection that started it. Under the new one, it's the default.

## The wire format, the hard way

None of this is documented end to end yet; agentgateway's own release notes admit as much ("most of it is not yet covered by a dedicated guide"). I found the exact requirements by sending bare requests and reading agentgateway's error messages, which turned out to be precise enough to iterate against directly:

```
400 mcp: invalid MCP protocol version header
```

Fixed by adding `MCP-Protocol-Version: 2026-07-28` as an HTTP header, not just inside the JSON-RPC body.

```
400 invalid request parameters: _meta.protocolVersion is required for modern requests
```

Fixed by adding `_meta` to the JSON-RPC `params`, with the key spelled the reverse-DNS-qualified way, `io.modelcontextprotocol/protocolVersion`, not a bare `protocolVersion`.

```
400 invalid MCP routing header: Mcp-Method
```

Fixed by adding `Mcp-Method: tools/call` (or `tools/list`, matching the JSON-RPC method) as its own HTTP header, plus `Mcp-Name: cart_create` for tool calls. These are agentgateway's own routing headers, separate from the JSON-RPC envelope.

And trying to send a standard `initialize` call first gets:

```
{"jsonrpc":"2.0","id":1,"error":{"code":-32601,"message":"method not found: initialize"}}
```

With `statefulMode: stateless`, agentgateway rejects the handshake outright. That matches the spec's own framing: a modern, stateless client doesn't send one. At no point in any of this does agentgateway send or expect an `Mcp-Session-Id`, to the client or to the backend. `scripts/mcp_client.py` in the repo encodes all four requirements in one place so nobody else has to rediscover them by reading error strings.

## The proof

`make up` builds the demo image, creates a kind cluster, loads both images into it, and applies the namespace, the 3-replica Deployment, and agentgateway, waiting on each rollout. `make demo` runs `scripts/demo.py`, which does three things against the live cluster and asserts on the results. This is the unedited output from an actual clean run:

```
=== 1. Fan-out: 9 independent tools/call requests, no session affinity configured ===
  request 1: served by stateless-mcp-demo-757875b499-z2zsg
  request 2: served by stateless-mcp-demo-757875b499-w6mpr
  request 3: served by stateless-mcp-demo-757875b499-z2zsg
  request 4: served by stateless-mcp-demo-757875b499-w6mpr
  request 5: served by stateless-mcp-demo-757875b499-z2zsg
  request 6: served by stateless-mcp-demo-757875b499-w6mpr
  request 7: served by stateless-mcp-demo-757875b499-z2zsg
  request 8: served by stateless-mcp-demo-757875b499-z2zsg
  request 9: served by stateless-mcp-demo-757875b499-z2zsg
  -> 2 distinct replica(s) answered 9 requests: ['stateless-mcp-demo-757875b499-w6mpr', 'stateless-mcp-demo-757875b499-z2zsg']

=== 2. One logical session, four separate requests, zero shared server memory ===
  cart_create      -> pod=stateless-mcp-demo-757875b499-z2zsg      cartToken(len=108)
  cart_add_item    -> pod=stateless-mcp-demo-757875b499-w6mpr      items=['keyboard']
  cart_add_item    -> pod=stateless-mcp-demo-757875b499-w6mpr      items=['keyboard', 'mouse']
  cart_checkout    -> pod=stateless-mcp-demo-757875b499-w6mpr      total=$65.00 status=confirmed
  -> session correct across every hop, no Mcp-Session-Id, no sticky routing.

=== 3. Bonus: rolling restart of every replica, mid-session ===
  cart_create (before restart) -> pod=stateless-mcp-demo-757875b499-z2zsg
  triggered: kubectl rollout restart deployment/stateless-mcp-demo
  cart_add_item during rollout (1/5) -> pod=stateless-mcp-demo-757875b499-n9s5r
  cart_add_item during rollout (2/5) -> pod=stateless-mcp-demo-757875b499-z2zsg
  cart_add_item during rollout (3/5) -> pod=stateless-mcp-demo-7674b87667-9vvrw
  cart_add_item during rollout (4/5) -> pod=stateless-mcp-demo-7674b87667-9vvrw
  cart_add_item during rollout (5/5) -> pod=stateless-mcp-demo-7674b87667-rp9fm
  cart_checkout (after restart)  -> pod=stateless-mcp-demo-7674b87667-9vvrw total=$5.00 status=confirmed
  -> cart survived a full rolling restart: created on stateless-mcp-demo-757875b499-z2zsg, checked out on stateless-mcp-demo-7674b87667-9vvrw (that pod didn't exist when the cart was created).

All assertions passed.
```

Step 1 landed on 2 of the 3 replicas in this particular 9-request sample; a separate run during development hit all 3 of 3 in the first 6. The exact ratio isn't the point. The point is that no two consecutive calls were pinned together by any mechanism, and every call still succeeded. Step 3 is the one I'd reread twice: `cart_checkout` closes out a cart that was created on a pod from the Deployment's *previous* ReplicaSet, a pod that no longer exists by the time checkout runs. The replica that answers checkout has never seen that cart in its own memory, and decodes it correctly anyway, because the cart's state was never in that memory. Under the pre-2026-07-28 model, this exact sequence, no affinity plus a rolling restart mid-session, is close to the textbook way to break an MCP client.

## Correcting the premise

The angle I started with was that agentgateway shipped support for the new spec "the same day it landed." Checking `gh release list -R agentgateway/agentgateway` and `gh release view v1.4.0` shows that's not quite right, and the real timeline is more interesting. agentgateway v1.4.0, which added full 2026-07-28 support, published 2026-07-27 at 17:51 UTC, one day *before* the spec's dated revision went final. It was built against the spec's [release candidate](https://blog.modelcontextprotocol.io/posts/2026-07-28-release-candidate/), public since May 21, 2026, giving implementers over two months of lead time. v1.4.1 followed on 2026-07-29 with a round of MCP compatibility fixes, including validating upstream MCP responses against their expected type and no longer logging synthetic session identifiers as if they were real `Mcp-Session-Id` values. Willison's post and the Hacker News discussion both came a few days after that, on July 31 and August 1. "Day one" support is still accurate in spirit; "same day" was the wrong detail, and the actual sequence (built against the RC, shipped a day early, patched two days later) is a better story about how a spec revision this size actually gets implemented.

I've written before about federating MCP servers behind agentgateway for [multi-tenant deployments](/multi-tenant-mcp-federation/), where the axis of scale is different customers getting different tools from the same gateway. This post is a different axis: identical replicas of the same server, and the question is whether any of them can answer any request. Under the new spec, the answer is yes, with nothing extra configured to make it true.

The full repo, including the server, the Kubernetes manifests, the agentgateway config, and the scripts that produced the output above, is at [github.com/themsquared/stateless-mcp-scale](https://github.com/themsquared/stateless-mcp-scale). `make up && make demo` reproduces all three experiments on a kind cluster. The next thing worth testing is a multi-node cluster with a real HPA scaling the MCP server under load, to see the same no-affinity property hold when pods are moving across machines, not just across ReplicaSets on one node.

