AIDive

Video pack

MCP servers as the weak link: GhostSplice numbers, leak paths and a hardening checklist

11 min read

TL;DR

  • An MCP server is not a plugin, it is a credential holder that talks to your agent on every call: its tool descriptions and tool results land in the model context with the same weight as your own instructions.
  • GhostSplice shows that model alignment is not a defense: a theft order given in one block is refused 100% of the time by GPT-4o, Gemini 2.0 Flash and Llama 3.3; the same order split across a tool description and a tool result is obeyed 100% of the time.
  • The client matters as much as the model: Claude Haiku 4.5 refuses everything through the API and yields 100% in the three-fragment test run inside Cursor.
  • Supply chain is already a live wire: CVE-2025-6514 hit mcp-remote, an OAuth proxy downloaded more than 400,000 times, with a command injection triggered by a malicious server.
  • Cloudflare's protocol-level detection and WriteGuard are the first real enterprise controls, but they need Zero Trust with TLS inspection, and a local stdio server never shows up in them.
  • For a solo developer or a small team the defense is manual: inventory, provenance, scoped tokens, a human gate on every write, and treating tool output as data.

What the sources say

An MCP server bridges your agent and an external service, so it stores what it needs to connect in your place: tokens, API keys, service account credentials. The leak analysis published on August 17 starts from a blunt observation: tokens are pasted straight into configuration strings and stay readable on disk, one hurried commit away from a Git repository s2. The same analysis lists over-permissioning as the second hole: broad rights granted during development ship to production unchanged, so a single compromise exposes far more than actual usage justified s2. The fourth hole is prompt injection: an agent reads everything its tools bring back, a web page, a ticket, an internal document, and a hidden instruction inside that content gets followed as if it came from you, using legitimate tools to expose what they were meant to protect s2.

The third hole is the supply chain. CVE-2025-6514 affected mcp-remote, an OAuth proxy downloaded more than 400,000 times: a malicious server could trigger command injection on the developer's machine, run code and leave with the credentials. One popular npm package, installed in one line, and the door was open s3.

The ecosystem grew faster than its guardrails. The official registry passes 9,600 published servers, remote server deployments have multiplied by five since May 2025, anyone can publish, and there is no central validation; your agent treats a random entry with the same trust as an official tool s4. The NSA published a dedicated MCP security guide in May, stating that adoption of the protocol outpaced the construction of its protections s5.

GhostSplice, named by the ASSET research group, makes the agent perform the exfiltration itself. Instead of writing the full theft order, the malicious server splits it: one fragment in a tool description, the other in the result that tool returns. Each piece reads as harmless on its own; the agent recombines everything that enters its context and executes the whole instruction in good faith s1. The numbers are the point. With the instruction given in one block, GPT-4o, Gemini 2.0 Flash and Llama 3.3 refuse 100% of the time. With the fragmented instruction, all three comply 100% of the time s1. Claude models resist better on the surface, but Claude Haiku 4.5 refuses everything through the API and yields 100% in the three-fragment test run inside Cursor: the same model refuses in one client and exfiltrates in another, depending on the protections the client adds or does not add s1. What the tests stole: SSH keys, environment secrets, source code, customer data, on isolated projects with fake keys, with a published and reproducible method s1. The same lab published Ghostcommit in June, which hid its instructions in PNG files referenced by the project's conventions and then encoded the stolen secrets in source code as integers; instruction fragmentation is a family of attacks, not a one-off s1. Two prerequisites hold: the malicious server must already be connected to your agent, and the agent must have the right to read the targeted files s1.

On the enterprise side, Cloudflare calls the unapproved servers developers wire into their agents "shadow MCP". Since the spec update, every compliant MCP client sends an MCP-Protocol-Version header, and Gateway inspects that header on all analyzed TLS traffic, giving a security team a dashboard of unique servers, users and request volumes s6. The newest spec version adds Mcp-Method and Mcp-Name headers that expose the requested operation and the tool name without opening the request body, so the network can tell an agent reading a ticket from an agent deleting fifty. Cloudflare's rules cover two cases, pure shadow MCP (a server never approved) and portal bypass (an approved server reached directly), and both are blocked with the same base rule s6.

WriteGuard, opened in private beta, classifies every tool of every MCP server into a risk level and applies a different policy per level: a read passes without friction; a contained write such as posting a comment passes but is signed as coming from an agent on behalf of a named human, with an audit event sent to a central log; a critical action such as merging code, deploying to production or mass deletion is blocked before the server processes it s7. The GitLab example in the post: reading a merge request passes, commenting passes with attribution, merging is refused until a human does it. The agent keeps the permissions of the employee it serves, but each write carries two signatures, the person and the agent session. Cloudflare describes its own internal use: its portal connects 27 MCP servers, against 13 in April s7.

The limits are real. WriteGuard is private beta by sign-up, and Gateway detection requires a Cloudflare Zero Trust deployment with TLS inspection enabled s7. Detection only sees network traffic it decrypts: a local MCP server running over stdio as a plain process on your machine stays invisible to Gateway, and that is how most developer-installed servers run s6. None of these tools fixes the mechanism GhostSplice exposes. The ASSET researchers say the fix is to treat tool output as data, never as instructions, and that separation does not exist natively in agents yet. Their three recommendations: prevent a value output by one tool from feeding another tool's arguments unchecked, keep the ability to refuse each tool invocation by hand, and treat any annotation from an unverified server as hostile by default s1.

Verdict: what protects you and what does not

Control Who it serves Verdict
Model refusals Everyone Skip as a defense: 100% refusal in one block, 100% compliance when fragmented [s1]
Client-side protections Everyone Keep: the same model refused in the API and yielded in Cursor [s1]
Inventory and provenance of servers Solo and teams Keep: GhostSplice needs the server already connected [s1]
Dedicated scoped tokens, rotated Solo and teams Keep: plaintext tokens and over-permissioning are the first two leak paths [s2]
Pinning and auditing MCP dependencies Solo and teams Keep: mcp-remote shipped a command injection to 400,000+ downloads [s3]
Gateway header detection (MCP-Protocol-Version, Mcp-Method, Mcp-Name) Enterprises on Zero Trust Try if you already run TLS inspection; blind to stdio servers [s6]
WriteGuard risk levels Enterprises Try on the waitlist; private beta only [s7]
Human gate on every write, merge and delete Everyone Keep: the handmade version of what WriteGuard industrializes [s7]

Do this Monday

  • Dump the list of MCP servers actually connected to each of your agents and remove every one you have not used in the last month.
  • For each remaining server, write down who publishes it and read what it does with your data before keeping it; drop any server that came from a thread rather than the vendor.
  • Replace every shared or master credential in an MCP config with a dedicated token scoped to the minimum that server needs, and put a rotation date on it.
  • Check that none of your MCP config files are tracked in Git, and add them to .gitignore where they are not.
  • Pin the version of every MCP package you install, and check your lockfile for mcp-remote versions covered by CVE-2025-6514.
  • Turn on manual approval for any tool that writes, merges, deploys or deletes, and keep it on in every client you use.
  • Review tool descriptions of every third-party server once, looking for instructions aimed at the model rather than at you.
  • If you run Cloudflare Zero Trust, enable TLS inspection and build the shadow MCP dashboard from the MCP-Protocol-Version header.

Go further

  • Read the full GhostSplice write-up for the test matrix per model and per client, and for the three mitigations the researchers propose s1.
  • Look up Ghostcommit, the June attack from the same lab, to see how instructions hid in PNG files and how stolen secrets were encoded as integers in source code s1.
  • Work through the four leak paths of the August 17 analysis against your own configs: plaintext storage, over-permissioning, supply chain, prompt injection s2.
  • Read the NVD entry for CVE-2025-6514 and check which mcp-remote versions are affected before trusting any OAuth proxy in your stack s3.
  • Read the NSA design considerations for MCP: it is the one vendor-neutral checklist written for teams rolling out agent automation s5.
  • Study the MCP-Protocol-Version, Mcp-Method and Mcp-Name headers in Cloudflare's post even if you do not use Cloudflare: any proxy you control can log them s6.
  • Borrow WriteGuard's four risk levels (read-only, minimal impact, contained write, critical) as a review grid for the tools your own servers expose s7.
  • Browse the official registry README to understand what publication requires and what it does not check s4.

Sources

FAQ

Does using a safer model protect me from GhostSplice?

No. The attack never asks for anything forbidden in one piece, so refusal training does not trigger. The same Claude Haiku 4.5 refused everything through the API and complied 100% inside Cursor; the client's protections decided the outcome, not the model.

I am a solo developer, is any of the Cloudflare tooling for me?

Not today. Gateway detection needs a Zero Trust deployment with TLS inspection, WriteGuard is a private beta, and both are blind to local stdio servers. The checklist above is the solo version of the same controls.

Is a server listed in the official registry safe?

Listing is not validation. The registry passes 9,600 servers with no central review, and your agent trusts a registry entry as much as a vendor tool. Judge provenance and read the code, not the listing.

What is the single highest-value change?

Dedicated, minimally scoped tokens per server, rotated like production secrets. Plaintext master credentials in config files are the first leak path, and they make every other failure, from CVE-2025-6514 to prompt injection, far more expensive.