Skip to main content

We don't test our MCP profiles. We test the agent.

•
11 min read

When you connect an AI agent to a profile in MCP Profile Hub, you hand it real capability over your content, and it can read your models, fetch entries, and, with the right tools, change them. The question that decides whether that's safe isn't "is the API up?" It's "will the agent do the right thing with my content?" Here's how we make sure the answer is yes, and how we keep it yes on every change.

The short version

For an MCP server, the real user is an AI agent, and the real risk is the agent's behavior, not an HTTP status code.

  • We test MCP Profile Hub in two layers: a deterministic tool-functional matrix that proves every tool a profile advertises actually works, and agent-evals where a real model drives those tools and is graded on how it behaves.

  • The behavior layer grades far more than "did it answer": grounding, correct tool and argument use, multi-step chaining, multi-turn memory, consistency, an adversarial safety battery (jailbreaks and over-refusal), write boundaries, tool efficiency, and faithfulness.

  • We don't just run it once. Results are scored into a per-profile health number and checked against a committed baseline, so quality can't silently regress between releases.

  • The result on the latest run: every system profile healthy, zero failures, so you can build on MCP Profile Hub with confidence.

The reliability problem AI agents introduce

The Model Context Protocol lets an AI client (Claude, Cursor, and others) call a server's tools through a standard interface. MCP Profile Hub builds on it: teams curate profiles (focused toolsets) and expose them to their AI clients.

But an agent isn't a deterministic integration. It decides which tool to call, interprets the result, and composes an answer, probabilistically. So a server that is technically flawless can still be misused by an agent that hallucinates a number, calls the right tool with the wrong arguments, or is talked into a destructive action by a cleverly worded prompt. That's a genuinely new class of reliability risk, and it doesn't appear in the tests most teams already have.

Why "the API works" isn't "the agent behaves"

Our end-to-end suite already covers the conventional surfaces: the MCP Profile Hub dashboard flows and the protocol layer (connect, list tools, call tools, authentication, isolation, error contracts). Those are necessary, and they pass.

But a passing tools/list only proves a tool is advertised. It says nothing about whether the tool actually resolves when called, whether an agent will choose it, pass the right arguments, keep its answer anchored to your real data, or refuse when a prompt tries to make it delete something. Those are the properties that actually break for users. So we test them directly, in two layers.

Layer 1: a deterministic tool-functional matrix

Before we ask whether an agent uses a tool well, we prove the tool works at all. For every profile, we take its live tools/list and call each tool directly, resolving any required arguments from real data first (list content types, then fetch one; list entries, then fetch one), and assert three things: the call resolves and returns a result, a missing required argument is rejected with a proper error (not a 500), and a tool that's out of scope for the profile is refused, not executed.

No language model is involved here, so this layer is fast and essentially free, and it's the deterministic foundation the behavior layer stands on. When an agent-eval fails, this matrix is what tells you which thing broke: the tool, or the model's choice.

It earns its keep. On its first run it caught a real drift we'd otherwise have blamed on the model: a couple of use-case profiles were advertising tools that don't resolve when called (a stale seed vs. the live catalog). That's a server bug the behavior layer alone would have reported as "the agent failed," and we'd have chased the wrong thing.

Layer 2: a real agent, real tools, real data

On top of that foundation we connect a real model to each profile exactly the way a customer's AI client would (as an MCP connector over the Streamable HTTP transport, authenticated with a bearer token) and let it work against a live stack. Three choices make this a real behavioral test rather than a demo:

  • Nothing is mocked. The agent calls the real tools and gets real data back.

  • The model sees the full, untruncated tool result, just as a real IDE hands it over, with no trimming that would hide a problem (with only a generous safety cap to avoid a pathological payload).

  • Nothing the model returns is modified. Every assertion runs on the real bytes the agent produced.

Then we grade its answer against the evidence it gathered in that same run. That's what turns "did the agent do the right thing?" into something measurable.

Everything we grade

Cases are generated per profile from its live tools and its capability class: a read-only profile, a publishing profile, and the full-catalog cms profile each get the cases that make sense for them, and every profile gets the safety checks. (On a full run, a skill graded deeply on the first profile that exposes it runs shallow on the rest, so we don't re-pay for the same read twelve times.) Across those cases we measure:

able 1. What we grade, and what each check catches.

Check

What it verifies

Failure it catches

Grounding

Answer matches the agent's own tool results (real counts/names), and notes when results were paginated

Hallucinated or stale data

Correct tool use

The right tool is called for the request

Wrong or missing tool calls

Tool-correctness (argument-level)

On multi-step tasks, the agent fetched a discovered id, not a guessed one

Right tool, wrong (or invented) arguments

Multi-step chaining (up to 3 hops)

Discover a content type, list its entries, then fetch one entry's detail, graded on the outcome

Broken multi-step reasoning

Consistency

The same prompt run several times yields stable tool use and grounded results

Flaky, unreliable behavior

Multi-turn memory

Recalls a value from earlier in the conversation without re-fetching it

Context loss across turns

Adversarial safety battery

Refuses jailbreaks, indirect (tool-result) injection, credential-exfil, and hallucination bait

Unsafe, exploitable behavior

Over-refusal

A legitimate read phrased with urgency/pressure is still answered, not refused

False-positive refusals (over-blocking)

Write boundary

Read-only profiles never succeed a write tool

Privilege/scope violations

Out-of-scope

Unrelated asks are declined without fabricating or misusing tools

Confident nonsense

Completeness honesty

An answer never presents a partial (paginated) page as the whole set

A false sense of completeness

Tool efficiency (trajectory)

The answer is reached without duplicate/redundant tool calls

Wasteful or looping tool use

Faithfulness

The stricter of the judge score and grounding, i.e. min(judge, grounding)

Fluent fabrication a judge alone would pass

A few more signals ride alongside every case: no-truncation (the answer wasn't cut off at the token limit), latency (wall-time within budget), and a trajectory read — how many tool calls it took and whether any were redundant, so a leaner path scores as a more efficient agent. Prompts are also tested in paraphrased form so a profile can't pass just because it was asked in one exact phrasing. For the deepest runs we add opt-in checks: a multi-turn conversation eval (the agent must recall a value from an earlier turn without re-fetching it) and two metrics borrowed from the RAGAS/DeepEval playbook, claim-decomposition faithfulness (break the answer into claims, check each against the evidence) and task-completion (did the agent actually accomplish the goal). And where a profile can safely create, an opt-in write eval exercises the real create path end to end, on a disposable stack, cleaning up after itself.

How a case is scored, and why you can trust it

Grading an LLM is fuzzy, so we separate a hard gate from soft signals, make the gate resistant to noise, and refuse to let quality drift down quietly.

  • Hard gate (build-breaker): behavior + grounding. This is the only thing that fails the build. It runs with one retry to absorb LLM variance, so a single unlucky sample doesn't cause a false red.

  • Composite health score. A 0 to 1 blend of coverage (40%), accuracy (30%), format (20%), and faithfulness (10%), banded per profile as HEALTHY / REVIEW / ALERT.

  • Regression gate. Each run is compared against a committed, model-keyed baseline; if a profile's composite drops past a threshold, the run fails, so "still green" also means "no worse than the last known-good."

  • Blocked ≠ broken. If a profile can't be reached (auth/scope) or has no stack configured, the case is skipped, not failed, so environment gaps never masquerade as regressions.

Two details we're proud of: the judge is handed the agent's own tool evidence, so it grades against ground truth and can't wrongly claim the agent "made it up"; and faithfulness deliberately takes the minimum of the judge and grounding scores, so a fluent answer that isn't backed by data can't score its way to a pass.

The LLM judge is deliberately advisory — it never breaks the build on its own (that's the behavior + grounding gate's job). But we don't hide it when the judge disagrees with a passing case: the report surfaces those dissents prominently, each with a plain-language explainer of why it's flagged and why the case still passed. It's the honest middle ground: a second opinion stays visible for a human to weigh, without letting a fuzzy judge gate a release. The trajectory (tool-efficiency) and completeness signals are surfaced the same way — visible in the report, never a silent build-breaker.

Verified across the latest models

We test the profiles, not any one model, so we verify each profile against the latest frontier models rather than tuning to a single one. Every profile scored HEALTHY, which is the assurance that actually matters to you: a profile behaves the same regardless of which model your agent brings to it.

Because the checks are model-agnostic, the same evals run through any provider's native tool-use loop, and we can run several providers side by side in a single pass, so a profile's behavior is directly comparable model-to-model in one report. A tiered mode sends the hardest cases (multi-step chains and writes) to a stronger model. Baselines are model-keyed, so a HEALTHY result always means "healthy for this model," never an average that hides a weak one.

Verified continuously, not once

Trust that's true only at launch isn't trust. The deterministic matrix and the agent-evals both run in our continuous-integration pipeline, right after the functional suite, so every change to the Hub is re-verified (first that its tools resolve, then that agents behave) before it ships. If a change weakens grounding or safety, the hard gate catches it; if it merely erodes quality, the regression gate catches that too.

What this means when you build on MCP Profile Hub

The point isn't the test harness. It's the confidence it buys you:

  • The tools actually work. Before any agent touches a profile, the deterministic matrix has proven each advertised tool resolves, so an agent isn't handed a tool that 404s.

  • Answers you can rely on. Because grounding and faithfulness are enforced, an agent on an MCP Profile Hub profile is verified to speak from your actual content, not a plausible guess.

  • Safe by default. Tools are tested against a battery of jailbreak, injection, credential-exfil, and hallucination attempts, so exposing a profile to an AI client isn't an open door, and the same battery checks the agent doesn't over-refuse legitimate work either.

  • Boundaries respected. Read-only profiles are verified to decline writes, so the limits you configure are the limits the agent honors.

  • It stays good. Every change is re-graded and gated against a baseline, so the confidence you have today doesn't quietly erode tomorrow.

The robustness isn't a claim on a slide. It's a property we measure: deterministically and behaviorally, on every profile, on every change.

Our point of view

Here's the stance we've landed on, and the one we'd offer anyone building an MCP server: testing an MCP server means testing the agent's behavior, not just the protocol, and testing that the tools resolve underneath it. The protocol is the easy part. The hard, valuable part is whether a non-deterministic agent, handed your tools, stays grounded, safe, and predictable, and whether it stays that way as the product changes. That's the bar we hold MCP Profile Hub to, because that's the bar our customers' AI experiences are judged against.

Conclusion

AI agents are becoming first-class consumers of software, and content is one of the most consequential things you can let them touch. MCP Profile Hub is built and tested with that in mind: not just "does it respond," but "do the tools resolve" and "does the agent behave," verified continuously and gated against regression. That's what lets you connect your AI client and build with confidence.

Ready to try it? Head to the MCP Profile Hub documentation and connect your AI client to a profile.