Skip to content
anavem.com

ExplainerPublished 8 min read

Are MCP Servers Safe? Documented Risks and a Checklist

What the MCP specification and Anthropic docs say about MCP server risks, local vs remote, permissions and prompt injection, plus a checklist. No verdict.

By Emanuel DE ALMEIDA · Editor

In this article
  1. How this article was made
  2. What the sources say: the specification's principles
  3. What the sources say: local servers
  4. What the sources say: remote servers
  5. What the sources say: permissions in Claude
  6. What the sources say: directory labels
  7. What the sources say: prompt injection
  8. What to consider: a checklist before you connect a server
  9. Limitations, and when to avoid a server
Editorial evidence card for Are MCP Servers Safe? Documented Risks and a Checklist

Key takeaways

Documented
  • Answer: What the MCP specification and Anthropic docs say about MCP server risks, local vs remote, permissions and prompt injection, plus a checklist. No verdict.
  • Evidence: Based on 10 dated primary or official sources, most recently checked .
  • Scope: This article does not claim hands-on testing. Performance or safety verdicts require a linked test record.

No document can tell you that a given MCP server is safe or unsafe. The MCP specification says tools represent arbitrary code execution and must be treated with caution. This article collects the risks that official MCP and Anthropic pages document, compares local and remote servers, and gives a checklist of what to check yourself.

How this article was made

We read the pages listed in the sources on 2026-10-03. They are the MCP specification and security best-practices page (revision 2026-07-28), Anthropic's connector and Claude Code documentation, and one OWASP page. We did not test any server, Anavem has not audited any software, and we use no incident or vulnerability statistics, because we did not open an original report for any. Passages are labeled "What the sources say" and "What to consider" (our assessment, not a vendor statement).

The best-practices page is written mainly for people who implement MCP: authorization developers, server operators and security professionals. We use it to explain risks and to show what a well-built client or server is expected to do. You usually cannot check those implementation details yourself, which is why the checklist focuses on what you can see.

What the sources say: the specification's principles

The MCP specification (revision 2026-07-28) says MCP's capabilities come through arbitrary data access and code execution paths, which bring security and trust considerations. It lists three principles:

  • User consent and control. Users must explicitly consent to and understand all data access and operations, and keep control over what is shared and what actions are taken.
  • Data privacy. Hosts must get explicit consent before exposing user data to servers, and must not send resource data elsewhere without consent.
  • Tool safety. Tools represent arbitrary code execution. Descriptions of tool behavior, such as annotations, should be treated as untrusted unless they come from a trusted server. Hosts must get explicit consent before invoking any tool.

The page adds that MCP itself cannot enforce these principles at the protocol level. Applications and server authors are expected to build the consent flows.

The tools page says there should always be a human in the loop with the ability to deny tool invocations. It says clients should show tool inputs to the user before calling the server, to avoid malicious or accidental data exfiltration, and should log tool usage. Servers must validate inputs and apply access controls.

What the sources say: local servers

The security best-practices page describes local servers as servers running on your machine, which may have direct access to your system and may be reachable by other processes on it. It lists these risks when local servers have inadequate restrictions or come from untrusted sources:

  • arbitrary code execution with the client's privileges
  • no visibility into which commands run
  • command obfuscation that makes a malicious command look legitimate
  • data exfiltration
  • data loss, from attackers or from bugs in legitimate servers

For clients, the page says that one-click local setup must show the exact command without truncation, describe it as potentially dangerous, and require explicit approval. It says clients should warn that servers run with the same privileges as the client and should consider sandboxing. For server authors, it says stdio limits access to the client.

Anthropic's page on desktop extensions says an extension runs locally with your user permissions, can access only what you can access, and transmits no data unless it is designed to. The filesystem server README says its operations are restricted to the allowed directories you give it.

What the sources say: remote servers

Anthropic's page on connectors added by URL says a remote MCP server gives Claude tools that can read data from applications, create, modify or delete data, and take actions on your behalf. It lists these precautions: connect only to servers from trusted organizations, review requested permission scopes during sign-in, be aware of prompt injection, and monitor for unexpected changes in tool behavior.

Anthropic's connector verification page adds, for any connector not built by Anthropic, that the developer controls which tools it exposes and can change them at any time. Anthropic does not run a third-party connector's servers and does not control how it handles your data.

The best-practices page also covers server-side and sign-in attacks: confused deputy attacks, token passthrough, server-side request forgery, state handle hijacking, mix-up attacks and others. They matter to operators and client builders. One is relevant to users: the page says broad access scopes increase the impact of a stolen token, and recommends starting with a minimal scope set and raising it only when needed.

What the sources say: permissions in Claude

According to the connectors overview, you choose per conversation whether Claude can use a connector. Claude can ask your approval before using a tool: Allow once, or Always allow. Under the connector's tool permissions you can set a tool or group to Always allow, Needs approval or Blocked. The security section of the unlisted-connector page advises clicking Always allow only for trusted servers, turning off connectors you are not using, and blocking tools you do not need.

In Claude Code, the MCP page says to verify you trust each server before connecting it, and that servers fetching external content can expose you to prompt injection risk. Project-scoped servers in a .mcp.json file prompt for approval in interactive sessions. The security page says plugins can add servers too, so reviewing .mcp.json does not show every server a session can load. It also says Anthropic reviews connectors against listing criteria before adding them to its directory but does not security-audit or manage any MCP server.

What the sources say: directory labels

The verification page defines three labels. Verified means Anthropic tested the tools for quality and compatibility and the connector met its Software Directory Policy at the time of review. The page states that verification is not a security audit or a guarantee, and that the developer controls the tools, which can change after review. Community means a third party built it and Anthropic screened it but did not review it in depth. Custom means you added it yourself and Anthropic has not reviewed it.

What the sources say: prompt injection

Claude Code's security page defines prompt injection as a technique where an attacker tries to override or manipulate an AI assistant's instructions by inserting malicious text. OWASP's LLM01:2025 page separates direct injection from indirect injection, which occurs when a model accepts input from external sources such as websites or files. Its listed mitigations include least-privilege access and human approval for high-risk actions.

Anthropic's pages say Claude has built-in protections. Claude Code's security page also says no system is completely immune to all attacks.

What to consider: an MCP server that returns web pages, emails, tickets or documents can carry text written by someone else into the conversation. The tool you approved may not be the problem; the content it fetches can be.

What to consider: a checklist before you connect a server

This is our checklist, built from the sources above. It does not make a server safe, and Anavem has not audited any server.

  1. Who runs it? Find the publisher and whether the server runs on your machine or on their servers.
  2. What label does it show? Verified, Community or Custom. Remember that Verified is not an audit.
  3. What can its tools do? Read the tool names and descriptions. Separate read-only tools from tools that write, send or delete.
  4. What does sign-in ask for? Read each permission scope. Prefer the narrowest set.
  5. Which account are you connecting? Consider a separate account with limited access to the data you can lose.
  6. Where does your data go? Read the provider's privacy terms. Anthropic says it does not control a third party's data handling.
  7. What content will it pull in? Servers that fetch external content raise prompt injection exposure.
  8. Set approvals. Use Needs approval or Blocked for write tools. Use Always allow only for servers you trust.
  9. Local servers: read the exact command, restrict folders to one project folder, and do not run it from your home directory.
  10. Never paste secrets such as passwords or API keys into a chat to configure a server.
  11. Watch for change. Tools can change after you connect. Re-read the tool list and your approvals now and then, and disconnect what you no longer use.
  12. Report problems. Anthropic's page points to its bug bounty program for malicious MCP servers.

Limitations, and when to avoid a server

  • The sources are specifications and guidance, not measurements. They tell you what to look for, not how likely a problem is.
  • The implementer-side attacks cannot be checked by a user.
  • Documentation changes. Revision 2026-07-28 is the one we read.
  • If you cannot identify who runs a server, cannot read its tools before connecting, or cannot limit what it can reach, the checklist above suggests holding off. That is our assessment.

Related reading: what is an MCP server, are Claude Skills safe, the terminology map, and Anavem's MCP server index.

Sources

Get new guides by email

New verified tool profiles, tested workflows and pricing changes. Sponsored items are labelled.