MCProbe Whitepaper · v0.1 · October 2026

Vetting MCP servers before an agent trusts them

Model Context Protocol servers are becoming the package ecosystem of AI agents. People install them the way they installed npm packages in 2015: by URL, from a README, without reading what is inside. MCProbe is the check that should happen first.

1. The problem

An MCP server hands an agent three kinds of text it will act on: server instructions, tool descriptions, and tool results. Agents are built to follow text. That makes every one of those channels an instruction channel for whoever controls the server.

None of this is theoretical. Each item above is reproduced by the evil fixture server shipped with MCProbe, and the industry already has names for them: MITRE ATLAS now lists AI Agent Tool Poisoning (AML.T0110) and AI Supply Chain Rug Pull (AML.T0109) as techniques.

2. Manifesto

Treat every MCP server as untrusted input until it has been read, probed and graded. Nothing an agent connects to should be a black box.
  1. Read everything the agent will read. Instructions, every description, every schema, every prompt and resource.
  2. Try it before you trust it. Static review misses behaviour. Call the tools, safely, and inspect what comes back.
  3. Code decides what is safe to call, not the model. The model picks interesting tools; a deterministic gate refuses anything destructive and cannot be talked out of it.
  4. Every finding carries its evidence. No "looks suspicious". The offending text, the tool, the response.
  5. Speak to the person who has to decide. A grade, a sentence, and plain-language findings first; the JSON for engineers second.
  6. Map to the language the security world already uses. MITRE ATLAS, OWASP, ATT&CK, so a finding can go straight into a ticket or a risk register.

3. What MCProbe does

You give MCProbe an MCP server URL (Streamable HTTP, with SSE fallback) and optionally a bearer token. About twenty seconds later you get:

Nothing is installed on your machine. MCProbe speaks to the server from its own process.

4. How it works

An audit runs five phases. Every phase is isolated: a failure is recorded and the run continues to a report, so you always get an answer.

4.1 Connect

MCProbe opens a real MCP session with the official client SDK. Streamable HTTP first, SSE if that fails, 15 second timeout. It records transport, protocol version, server name and version, and the server's instructions. A server that cannot be reached, or answers but does not speak MCP, grades F with a single connectivity finding rather than an error page.

4.2 Discover

tools/list, prompts/list, resources/list, each tolerant of "method not supported". Each tool's name, description and schema are tokenised (cl100k) so the report can say exactly what the server costs an agent on every turn.

4.3 Analyze: deterministic checks

Pure code, no model, runs in milliseconds, cannot be prompt-injected.

CheckWhat it looks forSeverity
Agent-directed textPhrases in instructions or descriptions aimed at the agent: "ignore previous instructions", "do not tell the user", "before using this tool", <IMPORTANT>, "include the contents of", "send … to http"High for strong patterns, medium for weak ones
Per-tool context costA single tool definition over 400 tokensMedium
Total context costWhole inventory over 6,000 tokens, or more than 30 toolsHigh / medium
Dangerous capabilityexec, shell, eval, command, sudo, delete, remove, upload, drop, kill, credentials, passwords, tokens, secrets in a name or descriptionHigh if exec-class and unannotated, else medium
Schema qualityNo input schema, or a free-form object with no propertiesLow
TransportPlain HTTP to a non-local hostMedium

4.4 Analyze: LLM judge

One structured call. The model sees the instructions and the tool list and returns JSON: a risk level and an "is this addressed to the agent" flag per tool, plus a verdict on the instructions. It catches what regexes cannot: paraphrased manipulation, capability that does not match the server's stated purpose, deceptive descriptions. If the model answers with prose, MCProbe retries once, then proceeds without it. If the proxy rejects JSON mode, it falls back to plain mode. Default model: claude-haiku via the hackathon LiteLLM endpoint; every call is traced in neatlogs under the audit's session id.

4.5 Probe: a bounded agent behind a gate

This is what separates MCProbe from a linter. An agent loop with three functions, call_tool, record_finding, finish, is told to observe real behaviour with harmless inputs. It chooses which tools to call and with what arguments; code decides whether the call is allowed.

With MCPROBE_LLM=off the same phase runs without a model: every gate-approved tool is called once with placeholder arguments derived from its schema, with the same response checks and rug-pull diff. Lower coverage, zero cost, and it keeps the product working when the model is unavailable.

4.6 Report

Findings are deduplicated, scored, mapped to frameworks, persisted as JSON and streamed to the browser over Server-Sent Events as they happen. Scanning a URL that already has a report replays it instantly; "Re-scan live" forces a fresh run.

5. Threat model and framework mapping

MCProbe findingMITRE ATLASOWASP LLM Top 10 (2025)MITRE ATT&CK
Hidden instructions in a tool description (found by pattern, AI judge, or the probing agent)AML.T0110 AI Agent Tool PoisoningLLM01 Prompt InjectionT1195 Supply Chain Compromise
Tool definitions changed after useAML.T0109 AI Supply Chain Rug PullLLM03 Supply ChainT1195
Server instructions steer the agentAML.T0051 LLM Prompt Injection · AML.T0080 Agent Context PoisoningLLM01–
Tool output tries to hijack the agentAML.T0051 · AML.T0080LLM01–
Leaked credentialsAML.T0057 LLM Data Leakage · AML.T0098 Tool Credential HarvestingLLM02 Sensitive Information DisclosureT1552 Unsecured Credentials
Dangerous capabilityAML.T0053 AI Agent Tool Invocation · AML.T0101 Data Destruction via Tool InvocationLLM06 Excessive AgencyT1059 Command and Scripting Interpreter
Tries to send data elsewhereAML.T0086 Exfiltration via Agent Tool InvocationLLM02T1041 Exfiltration Over C2 Channel
Context bloat (tool list) and oversized tool responsesAML.T0034.002 Agentic Resource ConsumptionLLM10 Unbounded Consumption–
Truncated or malformed description–––
Impersonates another toolAML.T0110LLM01T1036 Masquerading
Unencrypted connection–LLM03T1557 Adversary-in-the-Middle

CVEs are deliberately absent. A CVE names a known bug in a specific software version; MCProbe finds behaviour in a server nobody has catalogued yet. The two are complementary: a future registry integration could attach CVEs to servers built on known-vulnerable packages.

6. Scoring

Start at 100. Subtract 30 per critical, 15 per high, 7 per medium, 2 per low, floor at zero. A ≥ 90, B ≥ 75, C ≥ 60, D ≥ 40, otherwise F. Weights are intentionally simple and visible; tuning them against a corpus of real servers is on the roadmap.

Servers that cannot be scanned get no grade. If MCProbe cannot complete an MCP handshake, the result is ? with a reason rather than an F: not_mcp (the address answered with a web page or something that is not MCP), auth_required (401 or 403; add a token and rescan), timeout, unreachable (DNS or connection refused), or blocked (a private or cloud-internal address MCProbe refuses to contact). An unknown is neither a pass nor a fail, and the report says so.

Keyword findings have three tiers. A dangerous word in the tool name, or an exec-class word anywhere (exec, shell, eval, command, sudo, kill, drop), keeps full severity. A word that appears only in the description is reported as low, "Possibly dangerous (keyword match)". A word the description explicitly negates ("read-only", "never returns", "does not") is not reported. An honest readOnlyHint or destructiveHint annotation steps the severity down one level; annotations never unlock the probe gate.

7. Status: what is implemented today

CapabilityStatusNotes
Streamable HTTP and SSE connection, bearer tokenImplemented15 s timeout, readable failure
Inventory of tools, prompts, resources with token costsImplemented
Deterministic checks (section 4.3)ImplementedSix checks; thresholds are constants; keyword rule has three tiers with negation handling
Unscannable servers reported as "?" with a reason, never gradedImplementednot_mcp, auth_required, timeout, unreachable, blocked
SSRF guard: private, loopback, link-local, CGNAT and cloud-metadata addresses refused, re-checked after redirectsImplementedMCPROBE_ALLOW_LOCAL=1 permits the local demo fixtures
Rug-pull diff: before/after of changed descriptions and schemasImplementedShown in the report with removed and added text
Per-finding impact and fix text, evidence highlighting, framework mapping per rowImplementedAlso in the JSON and CSV exports
Compare with previous scan, README badge, copy-ready safe-tools configImplemented
LLM judge with JSON output and fallbacksImplementedOne call per audit
Gated probing agent with caps, response checks, rug-pull diffImplemented5 calls, 8 turns
No-LLM modeImplementedMCPROBE_LLM=off
Scoring, grading, deduplicationImplemented
Framework mapping (ATLAS, OWASP, ATT&CK)ImplementedIDs verified against published sources
Live progress UI, verdict, findings table, exports, replayImplemented
Clean and evil fixture servers, 77 automated testsImplementedEvil grades F, clean grades A, end to end
Tracing of every model call (neatlogs), checkpoints (Entire)ImplementedOne session per audit
Hidden Unicode and homoglyph detection in descriptionsPlannedZero-width and bidi characters used to hide instructions
Tool-shadowing against a list of well-known tool namesPartialCategory and mapping exist; detector not yet wired
Prompt and resource content retrieval and inspectionPlannedToday only listed, not read
Per-tool judge batching with confidencePlannedReduces false positives on benign instructions
OAuth flows for authenticated serversPlannedBearer tokens only today
Local stdio servers and skill foldersPlannedSame checks, different transport
Per-audit LLM token budgetPartialCall and turn caps bound spend; no token ceiling yet
Threshold tuning against a corpus of public serversPlanned

8. Roadmap

Next: make the verdict trustworthy at scale

Then: fit into how teams work

Later: the agent gateway

The natural end state is MCProbe sitting between agents and servers at runtime: the same checks applied to live traffic, responses scrubbed of injected instructions before the model sees them, destructive calls requiring confirmation. The scanner is how that gateway earns the right to be trusted.

9. Limits and honesty

10. Business

Free scanner for individuals, which is also how the corpus is built. Paid tiers for teams: pin-and-watch monitoring per server, policy enforcement in CI, private registries, and the runtime gateway. The detailed pricing, cost and runway plan is being prepared with cfo.ai as part of the neatHack submission and will be linked here.

11. How it was built

MCProbe was designed and built in the 48 hours of neatHack 2026 by team SRP. Python 3.13, FastAPI, the official mcp client SDK, the hackathon LiteLLM endpoint for Claude models, a single HTML page for the UI. Every model call is traced with neatlogs, grouped per audit, which is how the before-and-after improvement (duplicate findings 5 → 0 on the same target) was measured. Commits are checkpointed with Entire. The design spec, implementation plan, and build log live in the repository.

MCProbe · neatHack 2026 · team SRP