# Building Workload Identity Federation for AI Agents on Kubernetes

> How an AI agent pod trades a Kubernetes service account token for a resource-scoped credential, with no shared secret and no long-lived key in the pod.

- Canonical URL: https://webofmike.com/workload-identity-federation-agents/
- Author: Mike Moore (https://webofmike.com/about/)
- Published: 2026-10-01
- Last modified: 2026-10-01
- Tags: Kubernetes, AI Agents, Security, MCP, Platform Engineering
- Cite as: Mike Moore, "Building Workload Identity Federation for AI Agents on Kubernetes", Web of Mike (webofmike.com), 2026-10-01. https://webofmike.com/workload-identity-federation-agents/


The easy way to give an agent pod access to an external API is a long-lived key in a Kubernetes Secret. It doesn't expire when the pod dies, it doesn't rotate itself, and if it leaks it grants everything it was ever scoped to grant, forever. I built [agent-wif-lab](https://github.com/themsquared/agent-wif-lab) to work through the alternative on a real cluster: an agent pod that never holds a credential at all, only a Kubernetes-issued token it exchanges, on every call, for something narrower.

This is workload identity federation, and it isn't new. AWS calls it IAM Roles for Service Accounts. GCP calls it Workload Identity Federation. Vault calls it the `kubernetes` auth method. All three hide the same mechanism behind a provider SDK call. This lab builds it from the primitives that mechanism is made of, so the shape is visible instead of assumed.

## The shape, before the cloud provider hides it

Every version of workload identity federation does the same three things in the same order:

1. The platform (Kubernetes) vouches for the workload's identity by issuing a signed, short-lived token.
2. A relying party verifies that token independently, against the platform's own public keys, with no shared secret.
3. The relying party mints a new token scoped to exactly the resource being requested, and hands it back.

`sts:AssumeRoleWithWebIdentity` does steps 2 and 3 inside AWS's STS service, using an OIDC federation you configured once. This lab does the same two steps in a service you can read the source of, called `wif-exchange`, that trusts nothing but the cluster's own signing keys.

```
agent pod                wif-exchange                  mock-cloud-api
   |  read projected SA token   |                              |
   |---------------------------->                              |
   |  POST /exchange?resource=  |                              |
   |     verify against         |                              |
   |     k8s /openid/v1/jwks    |                              |
   |     check roles.yaml       |                              |
   |<---- resource-scoped JWT --|                              |
   |                             |                              |
   |------------------- GET /data, Bearer <scoped JWT> -------->|
   |                             |    verify against            |
   |                             |    exchange's own JWKS       |
   |                             |    check aud == self         |
   |<--------------------------------------------- 200 data ----|
```

Three services, three trust boundaries, no shared secret anywhere in the diagram.

## Where the agent's identity actually comes from

Kubernetes has published an unauthenticated OIDC discovery document and JWKS endpoint on the API server since 1.24. `wif-exchange` reads it directly:

```python
def _fetch_cluster_jwks():
    token = _own_serviceaccount_token()
    headers = {"Authorization": f"Bearer {token}"}
    r = requests.get(f"{K8S_API}/openid/v1/jwks", headers=headers,
                      verify=CA_CERT_PATH, timeout=5)
    r.raise_for_status()
    return r.json()
```

That's the entire trust establishment. No client secret, no pre-shared key, no certificate Mike had to generate and distribute. The exchange service's own ClusterRole grants exactly the two `nonResourceURLs` it needs to read this, nothing else.

The agent pod's side of this is a projected service account token, not the default one Kubernetes mounts automatically:

```yaml
volumes:
  - name: wif-token
    projected:
      sources:
        - serviceAccountToken:
            path: token
            audience: wif-exchange
            expirationSeconds: 600
```

`audience: wif-exchange` matters as much as the 600-second expiry. A default service account token is scoped to `https://kubernetes.default.svc`, which authenticates the pod to the Kubernetes API and nothing else. This one is minted for a single named audience, so it can authenticate the pod to `wif-exchange` and nowhere else, even though both tokens come from the same signing key.

## Running the exchange end to end

The whole lab runs on a local `kind` cluster, no cloud account, no API keys:

```bash
kind create cluster --config kind-config.yaml
kubectl config use-context kind-agent-wif-lab
bash scripts/deploy.sh
bash scripts/run-demo.sh
```

`run-demo.sh` execs into the agent pod and runs four requests against the live cluster, and all four are real HTTP calls, not a recorded transcript:

1. Exchange the projected token for `/data`, a resource the service account's entry in `roles.yaml` authorizes → `200`, a resource-scoped access token comes back.
2. Call `mock-cloud-api`'s `/data` endpoint with that token → `200`, data.
3. Try to exchange the *same* projected token for `/billing`, which the service account is not authorized for → `403` from the exchange itself, before any token is even minted.
4. Replay the valid `/data` token against `/billing` anyway → `403` from `mock-cloud-api`, because the destination independently checks the audience rather than trusting that a validly-signed token is good for anything it's pointed at.

Step 4 is the one worth sitting with. A signature check alone answers "did the exchange service sign this," not "was this token meant for the thing being asked." `mock-cloud-api` checks both, and that's what makes replay against a different resource a 403 instead of a security incident.

The authorization model itself is one YAML file:

```yaml
"system:serviceaccount:agent-wif-lab:agent-sa":
  resources:
    - "https://mock-cloud-api.local/data"
```

No wildcard, no "any authenticated service account gets any resource." A verified Kubernetes identity maps to an explicit resource list, and the exchange refuses anything not on it.

## The audience check that also has to run backward

The main demo proves the exchange checks the audience on outgoing tokens. There's a fifth property, verified but deliberately left out of the main script so it doesn't muddy the four-step story: the exchange also has to reject an incoming token whose audience is wrong, not just whose signature is right.

```bash
kubectl -n agent-wif-lab exec agent-workload -- python3 -c "
import requests
with open('/var/run/secrets/kubernetes.io/serviceaccount/token') as f:
    tok = f.read().strip()
r = requests.post(
    'http://wif-exchange.agent-wif-lab.svc.cluster.local/exchange',
    params={'resource': 'https://mock-cloud-api.local/data'},
    headers={'Authorization': f'Bearer {tok}'},
)
print(r.status_code, r.json())
"
```

This uses the pod's *default* service account token, audience `https://kubernetes.default.svc`, mounted automatically whether you asked for it or not. Verified output:

```
401 {'error': "token verification failed: Audience doesn't match"}
```

Same signing key, same pod, same subject claim, wrong audience, and the exchange refuses it. Audience binding earns its keep exactly here: a token that authenticates a pod to the Kubernetes API server should not also authenticate it to an external identity broker, even though a naive verifier checking only the signature would happily accept it.

## Why this isn't SPIFFE, and when you'd want it to be

This lab uses Kubernetes' native projected service account tokens, not a SPIFFE/SPIRE workload API, because that's what a cluster has on day one with no extra control plane to run. I wrote about the SPIFFE version of agent identity in [SPIFFE identity for AI agents, end to end](/spiffe-identity-for-ai-agents/): a service mesh issuing short-lived X.509 or JWT SVIDs to every workload, which solves a related but distinct problem, mutual TLS between workloads inside the mesh. Workload identity federation, the pattern in this post, is about a workload reaching *out* to something that isn't in the mesh at all. Once east-west traffic between your own agents is the question, SPIFFE is the natural next layer on top of what's here, not a replacement for it.

## The gap this closes, and the one it doesn't

I wrote in [MCP Agent Identity: One Spec Shipped, Three Still Open](/mcp-agent-identity-gap/) that the MCP roadmap names Workload Identity Federation as a priority workstream, tracked as SEP-1933, and that it's still an open pull request. That post's conclusion was that the workload identity you need already exists a layer down, in the platform rather than the protocol. This lab is that layer down: nothing here is MCP-specific, which is exactly the point. An MCP tool call, an A2A message, or a plain REST call from an agent pod all cross the same Kubernetes-to-external-API boundary, and this pattern authenticates any of them without waiting for SEP-1933 to merge.

What it doesn't close: SEP-1933, when it ships, would let an *agent itself* carry a portable identity across MCP servers and clients, independent of which cluster or cloud it happens to be running in this week. The pattern here is cluster-scoped, verified against one cluster's JWKS. That's a real limitation for an agent that needs the same identity whether it's running in your cluster or a partner's, and it's the reason the spec work is happening at all.

## What's not covered

The lab makes no attempt at key rotation as a scenario. The exchange's signing key is generated in memory at pod startup; restarting the pod rotates it, which is realistic but not exercised as a demo step. And there's no real cloud provider here on purpose. The mapping onto AWS IRSA (`sts:AssumeRoleWithWebIdentity`), GCP Workload Identity Federation, or Vault's `kubernetes` auth method is direct: swap `wif-exchange` for the provider's own STS-equivalent endpoint, and the verify-then-exchange shape doesn't change.

Full source, kind cluster config, and the run scripts are in [themsquared/agent-wif-lab](https://github.com/themsquared/agent-wif-lab).

