Shadow MCP Servers: Why They're Hard to Track

Stop shadow MCP servers at an agentgateway chokepoint. Federate registered servers, authorize each tool call, then find what scans miss.

Tracking shadow MCP servers across teams is hard because most of them never show up anywhere a platform team looks. Local MCP servers run as child processes of an IDE or agent over stdio, so they have no network presence of their own. Their config is spread across a dozen files in different clients and scopes, some committed to repos. Remote servers sit behind one person’s OAuth grant, sometimes brokered from a vendor’s cloud. There’s no registry of record for internal servers, tool lists change at runtime, and the July 2026 spec removed the initialize handshake that older scanners depend on.

Defensive security research. This post analyses publicly disclosed findings so the controls that stop them can be tested. It describes mechanisms, not procedures, and links the primary sources for anyone verifying the work.

This page covers why tracking is structural, a scanner for the network half, a config-path table for the laptop half, and the controls that keep an inventory true. The scanner is at themsquared/shadow-mcp-scanner . The numbers on this page come from the repo’s local-process lab (scripts/lab_local.sh). A compose re-run is pending.

Why shadow MCP servers are hard to track across teams

OWASP lists this as MCP09:2025 Shadow MCP Servers . The difficulty is how the protocol is built, not that teams are careless.

Local servers have no network presence

One of the two standard transports is stdio : messages over the standard streams of a client-launched subprocess. The server never binds a port. Authorization is optional , and for stdio the spec says clients SHOULD NOT follow the HTTP flow. Credentials come from the environment.

What that means for tracking: there is no listener and no handshake on your network. The only packets are the server’s own calls to GitHub, Jira, or a database, which look like the developer’s normal traffic.

The config lives in a dozen files

There is no single MCP config file. Claude Code, Claude Desktop, Cursor, VS Code, Windsurf (now documented under Devin), Gemini CLI, Codex, and GitHub Copilot each have their own, split across user, project, and managed scopes. VS Code also reuses other tools’ configs .

A project .mcp.json arrives with git clone. Claude Code prompts for approval interactively, but in claude -p, the Agent SDK, and cloud sessions it loads those servers without asking. Windsurf’s current docs live under Devin and name ~/.config/devin/mcp_config.json .

What that means for tracking: one client’s user config is not an inventory. Miss a scope and you miss a server a repo is shipping to every clone. The table later in this page is the list to sweep.

Remote servers sit behind personal OAuth and vendor clouds

A remote MCP server a developer authorizes, such as a personal MCP connector like the Pocket AI recorder’s , is one person’s OAuth grant. Claude custom connectors go further: connections originate from Anthropic’s cloud, not the local device . Nothing crosses your network, there is no local config file, and there is no process to find.

What that means for tracking: a laptop sweep and a network scan can both be clean while a vendor-brokered connector is still live.

There’s no registry of record

The official MCP Registry is in preview and lists publicly accessible servers only. It does not support private servers such as mcp.acme-corp.internal, recommends you host your own registry that implements its OpenAPI spec, and says its codebase is not designed for self-hosting.

Name matching is not a substitute. GitHub Copilot’s registry policy matches on server name or ID and “can be bypassed by editing configuration files.” GitHub calls that preview, not the recommended method, and points at URL allowlists instead.

What that means for tracking: until you run a registry yourself, the two scans have nothing authoritative to reconcile against.

What a server can do changes at runtime

Tool lists can change while the server is running (listChanged, notifications/tools/list_changed, and in 2026-07-28 subscriptions/listen). List results carry ttlMs and cacheScope, so clients cache what they last saw. Clients MUST treat tool annotations as untrusted unless they come from trusted servers. Discovery can also carry a server-controlled instructions field that lands in the system prompt .

What that means for tracking: counting endpoints is not counting capabilities. If a description can change between calls, pin tool definitions, not just names .

The handshake changed in July 2026

The current protocol version is 2026-07-28 . That revision removed the initialize / notifications/initialized handshake and Mcp-Session-Id. Every server MUST implement server/discover . Legacy is 2025-11-25 and earlier. Modern is 2026-07-28 and later. Dual-era servers speak both.

What that means for tracking: older scanners miss modern-only servers. A legacy client hitting a modern-only HTTP server lacks the required headers and gets 400 Bad Request. Both eras are below.

How big the exposed half is

In July 2025, Alfredo Oliveira and David Fiser at Trend Micro published a count of internet-exposed MCP servers : 492 of them running with no client authentication and no transport encryption, collectively exposing 1,402 tools. More than 90% offered direct read access to a data source. About 74% were hosted on AWS, Azure, GCP, or Oracle.

That report is over a year old, and it is still the number people cite. Nobody has published a follow-up census. The interesting half is the server bound to a VPC, reachable by every workload in the cluster and every laptop on the VPN. That one is not in the 492, and it is the one your agents will actually find.

Fingerprinting an MCP server on the network, in both protocol eras

You cannot identify an MCP server by port. You identify it by asking a question only an MCP server answers. That question changed in July 2026.

Modern servers answer server/discover

The current fingerprint is a JSON-RPC server/discover POST. Streamable HTTP requires MCP-Protocol-Version and Mcp-Method on every request, so intermediaries can inspect without parsing the body . Version and client identity go in _meta:

POST /mcp
MCP-Protocol-Version: 2026-07-28
Mcp-Method: server/discover
Content-Type: application/json
{
  "jsonrpc": "2.0",
  "id": "discover-1",
  "method": "server/discover",
  "params": {
    "_meta": {
      "io.modelcontextprotocol/protocolVersion": "2026-07-28",
      "io.modelcontextprotocol/clientInfo": {
        "name": "ExampleClient",
        "version": "1.0.0"
      },
      "io.modelcontextprotocol/clientCapabilities": {}
    }
  }
}

A modern server replies with a DiscoverResult: supportedVersions, capabilities, optional instructions, and serverInfo in _meta. This shape is from the discover spec , not from the scanner in this post.

{
  "jsonrpc": "2.0",
  "id": "discover-1",
  "result": {
    "resultType": "complete",
    "supportedVersions": ["2026-07-28"],
    "capabilities": {
      "tools": {},
      "resources": {}
    },
    "_meta": {
      "io.modelcontextprotocol/serverInfo": {
        "name": "ExampleServer",
        "version": "1.0.0"
      }
    },
    "instructions": "This server provides weather and resource utilities.",
    "ttlMs": 3600000,
    "cacheScope": "public"
  }
}

serverInfo is self-reported. The spec says clients SHOULD NOT rely on it for security decisions. A name in that block is a label the server chose.

Legacy servers answer initialize

Servers on 2025-11-25 and earlier still speak initialize. Send it over HTTP POST. A real legacy server responds with its protocol version and a serverInfo block:

{
  "jsonrpc": "2.0",
  "id": 1,
  "result": {
    "protocolVersion": "2025-06-18",
    "capabilities": {
      "tools": {
        "listChanged": false
      }
    },
    "serverInfo": {
      "name": "mcp-open",
      "version": "1.4.2"
    }
  }
}

That handshake costs one request and needs no credentials. It is the right probe for servers that have not moved to 2026-07-28.

The deprecated HTTP+SSE transport

Before Streamable HTTP, MCP used HTTP+SSE. The changelog now lists that transport as Deprecated , eligible for removal. A GET /sse that comes back as text/event-stream with an event: endpoint line is still an MCP server, and easy to miss if you only speak the current transport.

A passive signal at the proxy

Modern Streamable HTTP also carries Mcp-Name on tools/call, resources/read, and prompts/get. A gateway or TLS-terminating egress proxy can label MCP traffic from those headers. Legacy traffic does not carry them. I have not pulled that header from my own logs. Treat it as a spec capability, not as a detection I have run.

The scanner now probes server/discover first, then initialize, then GET /sse. That order was verified against a modern-only lab server: mcp-modern reported OPEN / modern. If you are writing a probe today, use the same sequence.

Four postures, not two

The useful output is not a list of servers. It is a list of servers sorted by what they require of a caller. Both tables below are runs of the repo’s local-process lab (scripts/lab_local.sh), not docker compose.

The unmodified scanner only sent initialize. Pointed at the same seven-endpoint lab, it reported the 2026-07-28-only server (mcp-modern) as not MCP:

ENDPOINT                           POSTURE                    TRANSPORT            TOOLS
------------------------------------------------------------------------------------------------
http://legacy-sse:8080/sse         OPEN                       http+sse (legacy)    1
http://mcp-dual:8080/mcp           OPEN                       streamable-http      1
http://mcp-open:8080/mcp           OPEN                       streamable-http      3
http://mcp-bearer:8080/mcp         PROTECTED (no discovery)   streamable-http      -
http://mcp-oauth:8080/mcp          PROTECTED (RFC 9728)       streamable-http      -
http://build-dashboard:8080        not MCP                    -                    -
http://mcp-modern:8080             not MCP                    -                    -

scanned 7 endpoints: 5 speak MCP, 3 open with no auth

After the fix, the scanner sends server/discover first, then initialize, then GET /sse. The same lab reports mcp-modern as OPEN / modern, and the table gains a PROTOCOL column:

ENDPOINT                           POSTURE                    PROTOCOL     TRANSPORT            TOOLS
------------------------------------------------------------------------------------------------------------
http://legacy-sse:8080/sse         OPEN                       sse-legacy   http+sse (legacy)    1
http://mcp-dual:8080/mcp           OPEN                       dual         streamable-http      1
http://mcp-modern:8080/mcp         OPEN                       modern       streamable-http      1
http://mcp-open:8080/mcp           OPEN                       legacy       streamable-http      3
http://mcp-bearer:8080/mcp         PROTECTED (no discovery)   legacy       streamable-http      -
http://mcp-oauth:8080/mcp          PROTECTED (RFC 9728)       legacy       streamable-http      -
http://build-dashboard:8080        not MCP                    not-mcp      -                    -

scanned 7 endpoints: 6 speak MCP, 4 open with no auth

PROTOCOL is modern, legacy, dual, sse-legacy, or not-mcp. The scanner was not told which of those seven hosts were MCP servers. It was handed seven addresses and worked it out. build-dashboard is an ordinary web service, included so the output demonstrates the scanner is not simply flagging open ports.

Four of the six that speak MCP are open. One of them only speaks the deprecated HTTP+SSE transport, which is exactly the server a current-transport-only sweep would have missed. mcp-modern is the one an initialize-only sweep misses.

What an open server hands over

server/discover or initialize tells you a server exists. tools/list tells you what it can do, and an open server answers that from an anonymous caller as readily as from an authorized one:

  http://mcp-open:8080/mcp answered tools/list with no credentials:
    - run_query: Execute a read-only SQL statement against the analytics warehouse (warehous...
    - send_email: Send an email as noreply@ from the shared notifications mailbox.
    - list_customers: Return customer records from the CRM, including billing contact and plan.

The tool names in the demo are deliberately mundane and deliberately far-reaching, because that combination is the finding. Note what the truncation is hiding. The full description of run_query reads:

Execute a read-only SQL statement against the analytics warehouse (warehouse-prod.internal:5439, service account svc_analytics).

An internal hostname, a port, and a service account name, disclosed to anyone who can route to the endpoint, before a single tool is invoked. Tool descriptions are written for a model, in prose, by whoever built the server. They are not treated as a public surface, and on an open server that is precisely what they are.

This is the reason discovery has to come before the enumeration is anyone’s problem. The disclosure happens at tools/list, and tools/list does not need a credential on a server that never asked for one.

The 401 that tells you nothing

The two protected servers in the table both return a 401. Only one of them is useful to a client that finds it:

mcp-bearer: HTTP 401
  WWW-Authenticate: Bearer
mcp-oauth: HTTP 401
  WWW-Authenticate: Bearer resource_metadata="http://mcp-oauth:8080/.well-known/oauth-protected-resource"

The second one names a metadata document. Fetch it and you get the authorization server, the supported scopes, and how to present the token:

{
  "resource": "http://mcp-oauth:8080",
  "authorization_servers": ["https://idp.internal/realms/agents"],
  "bearer_methods_supported": ["header"],
  "scopes_supported": ["mcp:tools:read", "mcp:tools:invoke"]
}

This is not a style preference between two valid options. The MCP authorization spec says servers MUST implement RFC 9728 Protected Resource Metadata, and MUST use WWW-Authenticate on a 401 to indicate the resource metadata URL. That MUST is unchanged in 2026-07-28. A bare Bearer is protected and non-conforming. Whoever finds that server knows they need a credential and has no supported way to learn which authorization server issues it.

So the scanner reports those separately. PROTECTED (no discovery) is not a finding on the same severity as OPEN, but it is a finding: it marks a server that will be hard to onboard properly and was probably configured by hand. I wrote about the wider version of this gap in the MCP agent identity gap , where the roadmap names four identity workstreams and the ecosystem ships one.

Finding the servers a network scan can’t see

How can I see what MCP servers are running on employee machines? Read the config files, not the process list. Each MCP client stores its servers in known places: ~/.claude.json and a project’s .mcp.json for Claude Code, claude_desktop_config.json for Claude Desktop, ~/.cursor/mcp.json, .vscode/mcp.json or the VS Code profile’s mcp.json, ~/.gemini/settings.json, and ~/.codex/config.toml. Sweep those paths and your repos read-only through MDM, record client, scope, transport, and URL or command, and never copy the env or headers values, which hold secrets. Servers users connect through a vendor’s web app, such as Claude connectors, won’t be in any local file.

Where each client keeps its MCP config

Paths move. Re-check the vendor docs before you ship a sweep. Documented 2026-09-28:

ClientScopePathFormat key
Claude Codeuser / local~/.claude.jsonmcpServers
Claude Codeproject.mcp.json at the repo root (often committed)mcpServers
Claude CodemanagedmacOS: /Library/Application Support/ClaudeCode/managed-mcp.json; Linux/WSL: /etc/claude-code/; Windows: C:\Program Files\ClaudeCode\allowedMcpServers / deniedMcpServers
Claude DesktopusermacOS: ~/Library/Application Support/Claude/claude_desktop_config.json; Windows: %APPDATA%\Claude\claude_desktop_config.jsonmcpServers
Cursorglobal~/.cursor/mcp.jsonmcpServers
Cursorproject.cursor/mcp.jsonmcpServers
VS Codeworkspace.vscode/mcp.jsonservers
VS Codeworkspace.mcp.jsonmcpServers
VS Codeuser profileprofile mcp.json (one per profile)mcpServers
VS Codedev containerdevcontainer.json → customizations.vscode.mcpMCP config block
VS Code / CopilotAgent Host~/.copilot/mcp-config.jsonCopilot MCP config
Windsurf (Devin Cascade)usermacOS/Linux: ~/.config/devin/mcp_config.json; Windows: %APPDATA%\devin\mcp_config.jsonMCP config
Gemini CLIuser~/.gemini/settings.jsonmcpServers
Gemini CLIproject.gemini/settings.jsonmcpServers
OpenAI Codexuser~/.codex/config.toml[mcp_servers.<name>]
Claude.ai custom connectorsaccountno local file (connection originates from Anthropic’s cloud)n/a
Cursor team / marketplace serversteamno local file unless someone also wrote them into mcp.jsonn/a

Docs: Claude Code , managed MCP , Claude Desktop , Cursor , VS Code , Windsurf / Devin , Gemini CLI , Codex , Claude connectors .

A read-only config sweep

A process list will not do this job. stdio servers are children of a trusted IDE, and a remote server may not be running when you look.

Walk the paths in the table, plus repos, for .mcp.json, .vscode/mcp.json, .cursor/mcp.json, and .gemini/settings.json. Parse JSON (mcpServers, servers, projects.*.mcpServers) and TOML (mcp_servers). Emit client, scope, file, name, transport, and the URL host or command basename. Never print env or headers values. Those fields hold API keys and tokens. The repo’s scanner/discover_local.py does that sweep: it is read-only, redacts env and headers, and its fixture tests pass.

Deploy the sweep through MDM. mcp-audit is a third-party walker; I have not run it and I am not endorsing it. Keep any walker read-only and redacted, including rows that come back empty because the server never lived in a file.

Running it

git clone https://github.com/themsquared/shadow-mcp-scanner
cd shadow-mcp-scanner
bash scripts/demo.sh

That starts the lab and scans it. The tables on this page are from scripts/lab_local.sh, not docker compose. A compose re-run is pending. To check the classification against what the repo claims:

bash scripts/verify.sh

Against something real, the scanner takes hosts and ports rather than a CIDR. Feed it from your own inventory, your mesh, or nmap output:

python3 scanner/scan.py mcp.internal:8080
python3 scanner/scan.py buildbox.internal:8000-8100
python3 scanner/scan.py a.internal:8080,b.internal:3000 --json

The form worth putting in a pipeline is the gate:

python3 scanner/scan.py mcp.internal:8080 --fail-on-open

It exits 1 when any endpoint answered server/discover or initialize without credentials and 0 otherwise. The scan is cheap enough to run on every deploy, and the failure is unambiguous.

The scanner is read-only. It calls server/discover or initialize, then tools/list, never tools/call. It is standard library Python with no dependencies, which matters mostly because it means you can read all of it before you run it on your own network. Only scan networks you are responsible for.

What finding them does not fix

Discovery is the control that has to come first, and it is not the control that solves anything. A scan produces an inventory with a timestamp on it, and the inventory is stale the next time somebody runs docker compose up.

The durable version is a chokepoint. Route MCP traffic through a gateway that authorizes each tool call, so an unregistered server is not reachable by an agent even when it is running, and keep a registry that records who owns each server and what it is allowed to reach. I built the federation half of that in multi-tenant MCP federation (Solo Enterprise). The broader question of which controls actually bind an agent, versus which ones only look like they do, is the subject of the rogue agent control map .

How to stop shadow MCP servers

What actually stops a shadow MCP server from being used is enforcement at a point agents cannot route around, not a better scan. Put agentgateway in the only MCP path, standalone or on Kubernetes . Federate the servers you registered. Authorize each tool call with CEL mcpAuthorization . An unregistered server can still be running. The agent cannot reach it if every MCP hop goes through that gateway.

Put a gateway in the only path

Standalone MCP authorization runs CEL mcpAuthorization against tools/list and tools/call and filters disallowed tools from the list. Kubernetes tool access is the same control. Support for the 2026-07-28 revision landed in agentgateway v1.4 . Current releases can put dual-era servers behind one federated endpoint.

Narrow what a registered server exposes with an MCP method allowlist at the gateway . Test the policy with requests that should be denied. CEL authorization policies can fail open silently (tested on Solo Enterprise): both forms report healthy, and only a call you expected to refuse tells you which one you wrote.

Make egress the thing that enforces it

A gateway nobody has to use is a suggestion. Default-deny egress from agent workloads so remote MCP endpoints are reachable only through that gateway. Decide on the address you resolved, not the hostname the workload typed. That is egress decided at CONNECT time . That demo used a Python CONNECT stand-in to show the pattern, not agentgateway.

Laptops are the gap. A developer on a machine with unrestricted egress can still reach a personal remote server, and no cluster NetworkPolicy will notice.

Allowlist by URL and command, never by name

A name is a label. Claude Code’s managed MCP controls key allowedMcpServers and deniedMcpServers on serverUrl or serverCommand, and say “a serverName entry, in either list, is not a security control.” Cursor’s enterprise MCP Allowlist keys on command pattern for stdio and URL pattern for remote servers. Copilot’s name-matching policy is the same bypass, which is why GitHub points at URL allowlists instead.

If the allowlist is friendly names, anyone who can edit a config file can pick one you already approved.

Keep a registry of record you run yourself

The official registry will not do this. Run a private one that implements the official OpenAPI, and record owner, domain, and what each server can reach. The useful move is to register servers by business domain , not by whoever stood the process up. That registry is what the config sweep and the network scan reconcile against.

Bind servers to identities, not labels

A name, a serverInfo block, or a registry ID is a claim. Bind the gateway route to a workload identity and to audience-bound tokens, so “is this our server” is a verified credential. SPIFFE workload identity is the attested half of that. The rest of the argument sits in the agent identity on Kubernetes series .

Until that exists, the scanner tells you the size of the problem. Four open servers out of seven is a demo number. The one worth knowing is yours.

The tables and commands on this page were run against the repo’s local-process lab (scripts/lab_local.sh) on the commit that ships in the repo. A compose re-run is pending.

Frequently asked questions

How do I stop shadow MCP servers?

Put agentgateway in the only MCP path, standalone or on Kubernetes. Federate the servers you registered and authorize each tool call with CEL mcpAuthorization, so an unregistered server is not reachable by an agent even while it is running. Default-deny egress so remote MCP is reachable only through that gateway. Allowlist by URL or command, never by name. Keep a registry of record you run yourself.

How do I find MCP servers running on my network?

Send a JSON-RPC server/discover POST with the MCP-Protocol-Version and Mcp-Method headers (the current spec). Fall back to initialize for older servers, then probe GET /sse for the deprecated HTTP+SSE transport. A reply identifies an MCP server regardless of port. Both current-spec and legacy checks are cheap, need no credentials, and are the only reliable way to tell an MCP endpoint from any other HTTP service.

What does an unauthenticated MCP server expose to a stranger?

Everything it can do. A tools/list call against an open server returns the name, description, and input schema of every tool. Descriptions routinely carry internal detail: in the demo, a run_query tool discloses the warehouse hostname warehouse-prod.internal:5439 and the service account name to an anonymous caller, before anything is invoked. Trend Micro found 492 such servers exposing 1,402 tools, over 90% with direct read access to a data source.

Does the MCP spec require OAuth discovery metadata on a 401?

Yes. The MCP authorization spec (2025-06-18 and the current 2026-07-28) says MCP servers MUST implement RFC 9728 Protected Resource Metadata and MUST use the WWW-Authenticate header on a 401 to indicate the resource metadata URL. A server that returns a bare WWW-Authenticate: Bearer is protected but non-conforming, and a client that finds it has no way to discover which authorization server issues its tokens.

Why is it so hard to track shadow MCP servers across teams?

Because most of them never show up anywhere a platform team looks. Local MCP servers run as child processes of an IDE or agent over stdio, so they have no network presence of their own. Their config is spread across a dozen files in different clients and scopes, some committed to repos. Remote servers sit behind one person's OAuth grant, sometimes brokered from a vendor's cloud. There's no registry of record for internal servers, tool lists change at runtime, and the July 2026 spec removed the initialize handshake that older scanners depend on.

How can I see what MCP servers are running on employee machines?

Read the config files, not the process list. Each MCP client stores its servers in known places: ~/.claude.json and a project's .mcp.json for Claude Code, claude_desktop_config.json for Claude Desktop, ~/.cursor/mcp.json, .vscode/mcp.json or the VS Code profile's mcp.json, ~/.gemini/settings.json, and ~/.codex/config.toml. Sweep those paths and your repos read-only through MDM, record client, scope, transport, and URL or command, and never copy the env or headers values, which hold secrets. Servers users connect through a vendor's web app, such as Claude connectors, won't be in any local file.

Does the official MCP Registry list internal MCP servers?

No. The official MCP Registry is in preview and lists only publicly accessible servers. It does not support private servers such as an internal hostname. For those, MCP recommends running your own registry that implements its API. The registry codebase is not designed for self-hosting.