Tags
The CopilotKit wordmark in white beside a blue and purple kite with a wavy blue tail, on a black card

Deterministic AI agent tests with AIMock — mocking LLMs, vector DBs, MCP, and A2A on one port

AIMock mocks the whole AI stack on one port (LLM, vector DB, MCP, A2A), so an agent test suite runs in under three seconds with identical output every time.

On this page

Introduction

Testing an application built on an LLM has an awkward problem: send the same prompt twice and you get two slightly different responses. A test that passed on Monday fails on Tuesday. The cause is not a bug in the code. It is the non-determinism of the model.

A 2026 agent application makes this worse. A single request touches six or seven services: not just the LLM but a vector database for retrieval-augmented generation (RAG), Model Context Protocol (MCP) tools, and inter-agent traffic over the Agent2Agent (A2A) protocol. How to mock each of them becomes the real question.

AIMock, the open-source mock server CopilotKit released in April 2026, answers all of them on one port. The test suite this article builds runs 18 cases in 2.35 seconds against mocked LLM, vector DB, and A2A services, with byte-identical output every run.

This article walks through the demo agent and its tests, recording and replaying real fixtures, LLMock’s chaos-injection API, and then the other half of the exercise: migrating Google’s codelab burger agent from a2a-sdk v0.2 to v1.0, deploying it to Cloud Run, and pointing the same TypeScript client at it.

What AIMock is

AIMock is an npm package that mocks every service an AI application talks to.

  • 11+ LLM providers (OpenAI, Claude, Gemini, Bedrock, Azure, Vertex AI, Ollama, Cohere, and others)
  • Full MCP JSON-RPC 2.0 support
  • A2A agent card discovery and server-sent events (SSE) streaming
  • Vector database mocks (Pinecone, Qdrant, ChromaDB)
  • Chaos testing: injecting 500s, malformed JSON, and mid-stream disconnects, both from a hand-rolled server and through the setChaos() API
  • Zero dependencies; Node.js built-ins only

The concept is one config file and one port for the entire AI stack.

Where it sits among similar tools

I went looking for competitors and found that this niche has almost no direct ones.

Tool Focus Overlap with AIMock
Keploy Records real traffic and generates tests. General-purpose APIs No support for AI-specific protocols (MCP, A2A, streaming SSE)
VCR.py / Polly.js Classic HTTP record and replay Usable for LLM mocking, but no AI-specific features such as tool calls
Portkey / Helicone LLM gateways with caching and replay Production-oriented. No fault injection, no MCP mocking

If you want to test an agent application including MCP and A2A, AIMock is currently the only option. If all you need is a mock of the OpenAI API, a VCR-style HTTP replay is enough.

Recording and replaying fixtures, and why it matters

Record and replay is one of AIMock’s core features: capture a real API response, then serve it back verbatim during tests.

Seeing it the first time, I wondered how it differed from retrying a middle step of a workflow the way n8n does. The problems turn out to be fundamentally different.

What record and replay solves:

# Monday: test passes
assert response.includes("SQL injection risk") ✅
# Tuesday: same code, same prompt, OpenAI rephrased
assert response.includes("SQL injection risk") ❌
# it returned "potential SQL vulnerability"

With fixtures, an LLM call behaves like a pure function. Same input, same output, every time.

Aspect Without fixtures With fixtures
Cost A real API call per test run Zero
Speed 2–30s per LLM call Under 1ms
Determinism Different output every time Byte-identical
CI stability Flaky on network, rate limits, API changes Stable
Edge cases Not reproducible Captured permanently

Is mocking in CI a cop-out?

It is a fair question whether a mocked CI can catch production problems. The answer is that it tests something different.

What you are testing Needs the real API?
Does the JSON parser handle this response shape? No, a fixture is enough
Does the retry logic work on a 500? No, chaos injection is better
Does the prompt produce good output? Yes, impossible against a mock
Does the UI render the response correctly? No, a fixture is enough

In practice the best pattern is a two-tier CI pipeline:

  • Every PR: fast mocked tests that verify code logic
  • Nightly or weekly: integration tests against the real API, for drift detection and prompt validation

AIMock’s drift detection complements this. Run daily in CI, it compares the SDK type definitions, the real API response, and AIMock’s own output three ways, catching a provider’s API change before your users do.

Building the demo: a smart purchasing concierge

To exercise the main features I wrote tests for a purchasing agent.

Architecture

One agent, four calls, three mocks on localhostA buyer agent card in the centre lists four numbered steps: retrieve, plan, ask the seller agent, evaluate. Step 1 calls VectorMock and step 3 calls a hand-rolled seller agent mock on the left; steps 2 and 4 both call LLMock on the right. VectorMock and LLMock come from AIMock, the seller agent mock is written by hand. The agent then returns a recommendation to the user.User: “a lightweight laptop for travel, budget $1,200”Buyer agent (TypeScript)1RetrievePast reviews and product knowledge2PlanLLM drafts a search strategy3Ask the seller agentA2A request for matching offers4EvaluateLLM ranks the offers that come backVectorMockAIMock · Pinecone-compatibleSeller agent mockHand-rolled · A2A v1.0LLMockAIMock · OpenAI /v1Recommendation returned to the user

One request leaves the agent four times. Every call lands on a mock running on localhost.

The tests use three things:

  • LLMock: mocks the OpenAI API (the LLMock class from @copilotkit/aimock)
  • VectorMock: a Pinecone-compatible vector DB mock (the VectorMock class)
  • An A2A mock: a seller agent I wrote by hand with http.createServer

Project layout

aimock-a2a-demo/
├── src/
│ ├── buyer-agent/
│ │ ├── index.ts # the main BuyerAgent orchestrator
│ │ ├── a2a-client.ts # A2A protocol client
│ │ ├── knowledge-base.ts # vector DB / RAG layer
│ │ └── reasoner.ts # LLM reasoning
│ └── types/
│ ├── a2a.ts # A2A protocol types
│ └── product.ts # product types
├── tests/
│ ├── setup.ts # AIMock setup
│ ├── buyer-agent.test.ts # main test suite
│ └── chaos.test.ts # chaos tests
├── fixtures/llm/ # recorded LLM responses
├── aimock.json # AIMock config
└── package.json

Step 1: setting up LLMock

LLMock registers fixtures by pattern-matching on systemMessage or userMessage.

tests/setup.ts
import { LLMock, VectorMock } from "@copilotkit/aimock";
const llmock = new LLMock({ port: 0 }); // random port
// route fixtures by the content of the system prompt
llmock.on({ systemMessage: "evaluating product offers" }, {
content: JSON.stringify({
topPickIndex: 0,
reasoning: "The MacBook Air M4 offers the best combination...",
confidence: 0.92,
}),
});
llmock.on({ systemMessage: "purchasing concierge" }, {
content: JSON.stringify({
refinedQuery: "ultrabook laptop under 1.3kg...",
priorities: ["weight under 1.3kg", "battery life 15+ hours"],
priceRange: { min: 900, max: 1300 },
}),
});
await llmock.start();
process.env.OPENAI_BASE_URL = `${llmock.url}/v1`;
process.env.OPENAI_API_KEY = "mock-key";

What I learned here: onMessage() matches only the text of the user message. When an application calls the LLM more than once (twice here, for planning the search strategy and for evaluating offers), similar user messages can match an unintended fixture. Routing on on({ systemMessage: "..." }) by the content of the system prompt is reliable.

Registration order also matters. LLMock returns the first fixture that matches, so register the more specific patterns first and the catch-all (onMessage(/.*/)) last.

Step 2: setting up VectorMock

VectorMock exposes a Pinecone-compatible API.

tests/setup.ts
const vectorMock = new VectorMock({ port: 0 });
vectorMock.addCollection("product-knowledge", { dimension: 1536 });
// register static results — the same results come back every time
vectorMock.onQuery("product-knowledge", [
{
id: "review-macbook-air",
score: 0.94,
metadata: {
brand: "Apple",
model: "MacBook Air M4",
text: "MacBook Air M4: 1.24kg, 18hr battery life...",
},
},
// ... more entries
]);
await vectorMock.start();

What I learned here: the QueryResult type has no text field. Text data goes in metadata.text, and the application reads it back as m.metadata?.text.

VectorMock can in theory be mounted on the same port as LLMock via mount(), but the path prefix is not stripped, so starting it on its own port is what actually works.

Step 3: mocking the A2A protocol

A2A is the open inter-agent protocol from Google and the Linux Foundation (v1.2, 150+ participating organizations).

Where it genuinely earns its place is cross-organization agent communication — your agent talking to another company’s agent, say a payment processor or a logistics agent. Using A2A between sub-agents inside one application is overkill; a function call is enough.

In this demo A2A sits between the buyer agent and the seller agent. Those are meant to be agents from different organizations, so it is a legitimate use.

I wrote the A2A mock by hand with http.createServer, conforming to A2A v1.0:

tests/setup.ts
const server = createServer((req, res) => {
// Agent card discovery: GET /.well-known/agent-card.json (v1.0)
if (req.url === "/.well-known/agent-card.json" && req.method === "GET") {
res.writeHead(200, { "Content-Type": "application/json" });
res.end(JSON.stringify({
name: "Mock Electronics Seller",
skills: [{ id: "product-search", name: "Product Search", ... }],
...
}));
return;
}
// JSON-RPC: POST / — SendMessage (v1.0)
if (req.method === "POST") {
const rpc = JSON.parse(body);
if (rpc.method === "SendMessage") {
res.end(JSON.stringify({
jsonrpc: "2.0", id: rpc.id,
result: {
message: {
role: "ROLE_AGENT",
parts: [{ data: { offers: mockOffers } }],
messageId: `reply-${rpc.id}`,
contextId: rpc.params?.message?.contextId,
},
},
}));
}
}
});

Step 4: running the tests

Eighteen test cases, all finishing inside three seconds.

$ pnpm test
✓ tests/buyer-agent.test.ts (8 tests) 46ms
✓ A2A: Seller Agent Discovery > discovers seller agent card
✓ A2A: Seller Agent Discovery > checks skill availability
✓ A2A: Seller Agent Discovery > sends message and receives offers
✓ Vector DB: Knowledge Base > returns deterministic results
✓ Vector DB: Knowledge Base > same query twice returns identical results
✓ Vector DB: Knowledge Base > queryWithFallback returns empty on unreachable server
✓ Integration: Full Purchasing Flow > end-to-end: query → RAG → LLM → A2A → LLM
✓ Integration: Full Purchasing Flow > works without RAG knowledge
✓ tests/chaos.test.ts (10 tests) 2026ms
✓ Chaos: Vector DB failures > survives connection refused
✓ Chaos: Vector DB failures > respects timeout
✓ Chaos: Vector DB failures > throws on server error
✓ Chaos: LLM failures > malformed JSON
✓ Chaos: LLM failures > empty response
✓ Chaos: LLMock setChaos() > malformedRate — every response is garbled
✓ Chaos: LLMock setChaos() > dropRate — every request gets 500
✓ Chaos: LLMock setChaos() > clearChaos() restores normal operation
✓ Chaos: LLMock nextRequestError() > single request fails, next succeeds
✓ Resilience > agent degrades gracefully
Test Files 2 passed (2)
Tests 18 passed (18)
Duration 2.35s

Two of them are worth showing.

Proving determinism:

tests/buyer-agent.test.ts
test("same query twice returns identical results (determinism proof)", async () => {
const results1 = await kb.query("laptop");
const results2 = await kb.query("laptop");
// real Pinecone gives no such guarantee
expect(results1).toEqual(results2);
});

Graceful degradation:

tests/buyer-agent.test.ts
test("works without RAG knowledge (graceful degradation)", async () => {
const agent = new BuyerAgent({
sellerAgentUrl: sellerUrl,
knowledgeBase: {
apiKey: "mock-key",
baseUrl: "http://localhost:1", // unreachable → falls back
indexName: "product-knowledge",
},
ragTimeoutMs: 500,
});
const result = await agent.findProduct({
query: "I need a lightweight laptop for travel",
});
// even with RAG down, the agent still works via LLM + A2A
expect(result.topPick).toBeDefined();
expect(result.knowledgeUsed).toHaveLength(0);
});

That test is hard to write without mock infrastructure. Stopping a real Pinecone instance on demand is not an option.

The state of the A2A ecosystem

While researching the protocol I went looking for public agents you can freely interact with.

What exists

  • Hello World Agent: https://hello-world-gxfr.onrender.com/.well-known/agent.json. The first public A2A agent. Echo responses only, but useful for checking a client works
  • A2A Registry Playground: a2a-registry.org/playground, a playground for talking to registered agents
  • Google codelab: Purchasing Concierge, a hands-on that deploys two agents to Cloud Run and has them talk over A2A. Runs inside the free tier

An honest assessment

What exists What does not yet
Hello world / echo agents Production A2A endpoints published by real companies
The Google codelab demo A2A APIs from the likes of Stripe or Twilio
Community toy agents An agent marketplace with genuinely useful capabilities

A2A has 150+ participating organizations and over 22,000 GitHub stars, but there are almost no useful public agents you can freely interact with. Enterprise A2A adoption (Azure AI Foundry, Bedrock AgentCore) sits almost entirely behind authentication.

The protocol is mature (v1.2, under the Linux Foundation). The agents have not caught up.

Running a real A2A agent from the Google codelab locally

The first half of this used a mocked A2A seller agent written with http.createServer. The next step was to run a real A2A agent locally: the one from Google’s Purchasing Concierge codelab.

Cloning the repo and adjusting it

Terminal window
git clone https://github.com/alphinside/purchasing-concierge-intro-a2a-codelab-starter.git
cd purchasing-concierge-intro-a2a-codelab-starter/remote_seller_agents/burger_agent

This burger agent is an A2A seller agent built on CrewAI, LiteLLM, and Vertex AI Gemini. It takes burger orders from a user and processes them.

Four adjustments were needed against the original repo.

1. Move to Python 3.13

pyproject.toml
requires-python = ">=3.13"
Dockerfile
FROM python:3.13-slim

2. Regenerate uv.lock

A uv.lock hardcodes the full download URL of every package, so it carries whichever index the machine that generated it was pointed at. Mine had been generated behind a security proxy mirror, and every URL in the lock file went through it. In an environment that cannot reach that mirror, downloads time out.

Terminal window
rm uv.lock
UV_DEFAULT_INDEX=https://pypi.org/simple/ uv lock --python 3.13

This follows from uv sync --frozen using the URLs hardcoded in the lock file as-is. Changing the index configuration does not change the URLs already inside the lock file. It has to be regenerated.

Overriding the index inside the project too is worth doing:

pyproject.toml
[[tool.uv.index]]
url = "https://pypi.org/simple/"
default = true

3. Add litellm as an explicit dependency

agent.py does import litellm, but litellm was not in the pyproject.toml dependencies. It used to arrive as a transitive dependency of CrewAI; newer versions have dropped it.

Terminal window
UV_DEFAULT_INDEX=https://pypi.org/simple/ uv add litellm

Migrating a2a-sdk v0.2 to v1.0

The original code targets a2a-sdk v0.2, but the >=0.2.16 constraint installs v1.1.2, whose breaking changes produce import errors. I updated the code for v1.0.

The main changes:

Item v0.2 v1.0
Server construction A2AStarletteApplication class create_jsonrpc_routes() / create_agent_card_routes() route factories
AgentCard url field supported_interfaces list (AgentInterface)
RequestHandler no agent_card agent_card is required
Part construction Part(root=TextPart(text="...")) Part(text="...") (Protobuf-based)
JSON-RPC method tasks/send SendMessage
Agent card path /.well-known/agent.json /.well-known/agent-card.json
Protocol version none (implicitly 0.3) an A2A-Version: 1.0 header is required
Error type ServerError use UnsupportedOperationError directly

__main__.py, server construction:

__main__.py
from a2a.server.routes import create_agent_card_routes, create_jsonrpc_routes
from a2a.types import AgentCard, AgentInterface, AgentCapabilities, AgentSkill
from starlette.applications import Starlette
agent_card = AgentCard(
name="burger_seller_agent",
description="Helps with creating burger orders",
version="1.0.0",
capabilities=AgentCapabilities(streaming=True),
supported_interfaces=[
AgentInterface(protocol_binding="JSONRPC", url=agent_host_url)
],
skills=[skill],
)
request_handler = DefaultRequestHandler(
agent_executor=BurgerSellerAgentExecutor(),
task_store=InMemoryTaskStore(),
agent_card=agent_card, # required in v1.0
)
routes = []
routes.extend(create_agent_card_routes(agent_card))
routes.extend(create_jsonrpc_routes(request_handler, "/"))
app = Starlette(routes=routes)
uvicorn.run(app, host=host, port=port)

agent_executor.py, Protobuf-based types:

agent_executor.py
from a2a.types import Message, Part, Role, Task
from a2a.utils.errors import UnsupportedOperationError
class BurgerSellerAgentExecutor(AgentExecutor):
async def execute(self, context, event_queue):
query = context.message.parts[0].text if context.message else ""
result = await self.agent.invoke(query, context.context_id)
# v1.0: construct directly with Part(text="..."), no TextPart wrapper
await event_queue.enqueue_event(
Message(
role=Role.ROLE_AGENT,
parts=[Part(text=str(result))],
message_id=context.task_id or "msg",
context_id=context.context_id,
)
)

agent.py, going async:

CrewAI’s kickoff() is synchronous, but in a2a-sdk v1.0 the request handler is called from an async context. Calling a sync method inside an async context raises Agent execution was invoked synchronously from within a running event loop.

agent.py
# Before: def invoke(self, query, sessionId) -> str:
async def invoke(self, query, sessionId) -> str:
# ...
# Before: response = crew.kickoff(inputs)
response = await crew.kickoff_async(inputs)
return response

Confirming it works

Terminal window
# the Vertex AI API has to be enabled
gcloud services enable aiplatform.googleapis.com
# start it locally
uv run . --host 127.0.0.1 --port 8080

Checking the agent card:

Terminal window
$ curl http://127.0.0.1:8080/.well-known/agent-card.json | jq .name
"burger_seller_agent"

Sending a message over A2A v1.0:

Terminal window
$ curl -X POST http://127.0.0.1:8080/ \
-H "Content-Type: application/json" \
-H "A2A-Version: 1.0" \
-d '{"jsonrpc":"2.0","id":"1","method":"SendMessage",
"params":{"message":{"role":"ROLE_USER",
"parts":[{"text":"What burgers do you have?"}],
"messageId":"m1","contextId":"ctx1"}}}'

The response:

{
"result": {
"message": {
"role": "ROLE_AGENT",
"parts": [{
"text": "We have Classic Cheeseburger for IDR 85K, Double Cheeseburger for IDR 110K, Spicy Chicken Burger for IDR 80K, and Spicy Cajun Burger for IDR 85K. Do any of these sound good?"
}]
}
},
"id": "1",
"jsonrpc": "2.0"
}

The whole path works: Vertex AI Gemini → CrewAI → A2A v1.0.

Deploying to Cloud Run

With it working locally, the next step is Cloud Run. This burger agent is a Starlette ASGI server started by uvicorn, which makes Cloud Run the right fit rather than Cloud Functions. Cloud Functions invokes a Python function per request; an A2A agent is a resident HTTP server with several Starlette routes (/.well-known/agent-card.json and the JSON-RPC endpoint at /).

Terminal window
gcloud run deploy burger-agent \
--source . \
--port 8080 \
--allow-unauthenticated \
--region us-central1 \
--min-instances 0 \
--max-instances 1 \
--memory 1Gi

--source . runs a Dockerfile build on Cloud Build, stores the image in Artifact Registry, and deploys it to Cloud Run.

After deploying, set an environment variable so the agent card’s url field reflects the Cloud Run URL:

Terminal window
gcloud run services update burger-agent \
--region us-central1 \
--set-env-vars HOST_OVERRIDE=https://burger-agent-XXXXXXXXXX.us-central1.run.app

Checking it:

Terminal window
$ curl https://burger-agent-XXXXXXXXXX.us-central1.run.app/.well-known/agent-card.json | jq .name
"burger_seller_agent"
$ curl -X POST https://burger-agent-XXXXXXXXXX.us-central1.run.app/ \
-H "Content-Type: application/json" \
-H "A2A-Version: 1.0" \
-d '{"jsonrpc":"2.0","id":"1","method":"SendMessage",
"params":{"message":{"role":"ROLE_USER",
"parts":[{"text":"What burgers do you have?"}],
"messageId":"m1","contextId":"ctx1"}}}'
{
"result": {
"message": {
"role": "ROLE_AGENT",
"parts": [{
"text": "We have Classic Cheeseburger for IDR 85K, Double Cheeseburger for IDR 110K, Spicy Chicken Burger for IDR 80K, and Spicy Cajun Burger for IDR 85K. What would you like to order?"
}]
}
}
}

On cost: with --min-instances 0, instances scale to zero between requests and nothing is billed. The trade is a 10–15 second cold start, which is fine for a demo or test.

Bringing the TypeScript buyer agent to A2A v1.0

To connect to the real seller agent on Cloud Run, the TypeScript A2A client needed the v1.0 treatment too.

Changes to the A2A client

src/buyer-agent/a2a-client.ts
export class A2AClient {
// v1.0: agent-card.json (was: agent.json)
async discover(): Promise<AgentCard> {
const url = `${this.agentUrl}/.well-known/agent-card.json`;
const res = await fetch(url);
this.card = await res.json();
return this.card;
}
// v1.0: SendMessage plus the A2A-Version header
async sendMessage(text: string, contextId?: string): Promise<A2AMessage> {
const rpcRequest = {
jsonrpc: "2.0",
id: messageId,
method: "SendMessage", // was: "tasks/send"
params: {
message: {
role: "ROLE_USER", // was: "user"
parts: [{ text }], // was: { type: "text", text }
messageId,
contextId,
},
},
};
const res = await fetch(this.agentUrl, {
method: "POST",
headers: {
"Content-Type": "application/json",
"A2A-Version": "1.0", // required in v1.0
},
body: JSON.stringify(rpcRequest),
});
const rpcResponse = await res.json();
return rpcResponse.result.message; // was: result (a task object)
}
}

v0.3 to v1.0 changes, summarised

Item v0.3 v1.0
Agent card path /.well-known/agent.json /.well-known/agent-card.json
JSON-RPC method tasks/send SendMessage
Role values "user" / "agent" "ROLE_USER" / "ROLE_AGENT"
Part structure { type: "text", text: "..." } { text: "..." }
Header none A2A-Version: 1.0
Response result is a Task (array of messages) result.message is a single Message

Connecting to the real Cloud Run agent

I confirmed the TypeScript A2A client talking to the real burger agent on Cloud Run:

const client = new A2AClient("https://burger-agent-XXXXXXXXXX.us-central1.run.app");
const card = await client.discover();
// → Name: burger_seller_agent, Skills: ["create_burger_order"]
const reply = await client.sendMessage("What burgers do you have?");
// → role: ROLE_AGENT
// → "We have Classic Cheeseburger for IDR 85K, Double Cheeseburger for IDR 110K, ..."

The mocked tests and the real agent run through the same A2A client code and the same v1.0 protocol. Tests point at AIMock’s mock server; production points at the real Cloud Run agent. Only the URL changes.

Actually recording fixtures

Having covered record and replay in the abstract, I ran it.

Starting llmock

The @copilotkit/aimock package ships a CLI binary called llmock. The --record flag starts it as a recording proxy.

Terminal window
# start llmock in record mode, listening on port 4010
pnpm mock:record
# → llmock --record --provider-openai https://api.openai.com -f ./fixtures/llm

Running the app through llmock

Point OPENAI_BASE_URL at llmock and every OpenAI SDK call from the application goes through it. The application reads the API key from .env, and llmock forwards to OpenAI transparently.

Terminal window
OPENAI_BASE_URL=http://localhost:4010/v1 \
SELLER_URL=https://burger-agent-XXXXXXXXXX.us-central1.run.app \
pnpm dev

The recorded fixture

JSON like this lands in fixtures/llm/recorded/:

{
"fixtures": [{
"match": {
"userMessage": "User request: \"I need a lightweight laptop...\"",
"model": "gpt-4o-mini",
"turnIndex": 0,
"hasToolResult": false
},
"response": {
"content": "{\"refinedQuery\": \"lightweight travel laptop under $1200\", ...}"
},
"metadata": {
"systemHash": "a11524e7"
}
}]
}

The fields under match (userMessage, model, turnIndex) are the matching conditions. From then on, a request meeting the same conditions gets this response without OpenAI being called.

What it feels like in practice

The simplicity of the workflow is the surprising part:

  1. pnpm mock:record starts the proxy
  2. OPENAI_BASE_URL=http://localhost:4010/v1 pnpm dev runs the app normally
  3. Fixtures save themselves to fixtures/llm/recorded/
  4. From then on, pnpm mock:start replays them

No code changes at all. Only where OPENAI_BASE_URL points.

LLMock’s chaos injection

Four tests in that run (the setChaos() and nextRequestError() ones) have gone unexplained so far. They use LLMock’s programmatic chaos injection API, which reproduces LLM failure scenarios without the hand-rolled HTTP server the first five chaos tests stand up to fake a 500 or a timeout.

setChaos(): probabilistic fault injection

setChaos() sets a failure probability applied to every request.

tests/chaos.test.ts
test("setChaos({ malformedRate: 1 }) — every response is garbled", async () => {
const mock = new LLMock({ port: 0 });
// register a valid fixture (/.*/ matches every message)
mock.onMessage(/.*/, {
content: '{"refinedQuery": "test", "priorities": [], "priceRange": {"min": 0, "max": 100}}',
});
// chaos: corrupt the response 100% of the time
mock.setChaos({ malformedRate: 1.0 });
await mock.start();
process.env.OPENAI_BASE_URL = `http://localhost:${mock.port}/v1`;
const reasoner = new Reasoner("gpt-4o-mini");
// the fixture matches, but chaos breaks the response
await expect(
reasoner.buildSearchStrategy("laptop", [])
).rejects.toThrow();
await mock.stop();
});

ChaosConfig has three probability parameters:

Parameter Effect Works without a fixture
dropRate Returns a 500 Yes
malformedRate Returns {malformed json: <<<chaos>>>} No, a match is required
disconnectRate Cuts the TCP connection mid-stream No, a match is required

One gotcha: onMessage("*") matches the literal string *, not a wildcard. To match every message, use the regular expression /.*/.

clearChaos(): turning it off

clearChaos() removes the chaos configuration and normal fixture responses resume. Useful for testing recovery from a failure.

tests/chaos.test.ts
test("clearChaos() restores normal operation", async () => {
const mock = new LLMock({ port: 0 });
mock.onMessage(/.*/, { content: validJson });
mock.setChaos({ malformedRate: 1.0 });
await mock.start();
// ... env setup ...
const reasoner = new Reasoner("gpt-4o-mini");
// chaos on — it fails
await expect(
reasoner.buildSearchStrategy("laptop", [])
).rejects.toThrow();
// chaos off — it works
mock.clearChaos();
const strategy = await reasoner.buildSearchStrategy("laptop", []);
expect(strategy.refinedQuery).toBe("laptop"); // ✓ valid JSON
});

nextRequestError(): one-shot error injection

nextRequestError() returns an error for the next request only, then consumes itself. Ideal for a “fails once, recovers on retry” scenario.

tests/chaos.test.ts
test("nextRequestError(400) — single request fails, next succeeds", async () => {
const mock = new LLMock({ port: 0 });
mock.onMessage(/.*/, { content: validJson });
// one-shot: return 400 for the next request only
mock.nextRequestError(400, { message: "Bad Request" });
await mock.start();
// ... env setup ...
const reasoner = new Reasoner("gpt-4o-mini");
// first call: 400
await expect(
reasoner.buildSearchStrategy("laptop", [])
).rejects.toThrow();
// second call: the one-shot is spent → normal response
const strategy = await reasoner.buildSearchStrategy("laptop", []);
expect(strategy.refinedQuery).toBe("laptop"); // ✓
});

Note that the OpenAI SDK retries 5xx errors (500, 503, and so on) automatically. To test an error reliably with a one-shot, use a 4xx, which is not retried.

Test run

$ pnpm test
✓ tests/buyer-agent.test.ts (8 tests) 42ms
✓ tests/chaos.test.ts (10 tests) 1918ms
✓ Chaos: Vector DB failures > survives connection refused
✓ Chaos: Vector DB failures > respects timeout
✓ Chaos: Vector DB failures > throws on server error
✓ Chaos: LLM failures > malformed JSON
✓ Chaos: LLM failures > empty response
✓ Chaos: LLMock setChaos() > malformedRate — every response is garbled
✓ Chaos: LLMock setChaos() > dropRate — every request gets 500
✓ Chaos: LLMock setChaos() > clearChaos() restores normal operation
✓ Chaos: LLMock nextRequestError() > single request fails, next succeeds
✓ Resilience > agent degrades gracefully
Test Files 2 passed (2)
Tests 18 passed (18)
Duration 2.2s

Both approaches work — the hand-rolled HTTP server chaos tests (the first five) and LLMock’s built-in chaos API (the four after them; the last test, agent degrades gracefully, belongs to neither). The API version is less code and states its intent more clearly.

Summary

I built an AI agent test suite with AIMock, then migrated the real A2A agent from Google’s codelab to a2a-sdk v1.0, ran it locally, deployed it to Cloud Run, and connected to it from the TypeScript buyer agent.

Where AIMock earns its place:

  • Testing an agent application that combines an LLM, a vector DB, and several other services
  • Fast, cheap, stable test runs in CI/CD
  • Verifying fault tolerance through fault injection

Where it is overkill:

  • Mocking only the OpenAI API, where a VCR-style HTTP replay is enough
  • Evaluating prompt quality, which is impossible against a mock

Implementation notes for AIMock:

  • onMessage() matches only the user message. on({ systemMessage }) lets you route on the system prompt
  • onMessage("*") matches the literal *, not a wildcard. Use onMessage(/.*/) to match everything
  • Fixture registration order matters: the first match wins
  • QueryResult in VectorMock has no text field. Put it in metadata
  • Start VectorMock on its own port, separate from LLMock
  • malformedRate and disconnectRate in setChaos() need a fixture match. Only dropRate works without one
  • The OpenAI SDK auto-retries 5xx. Use a 4xx when testing errors with nextRequestError()

Lessons from the a2a-sdk v1.0 migration and deploy:

  • A2AStarletteApplication is gone. Use the route factories (create_jsonrpc_routes / create_agent_card_routes) inside a standard Starlette app
  • The type system moved from Pydantic to Protobuf. Construct directly, as in Part(text="...")
  • The agent card path changed from agent.json to agent-card.json
  • The JSON-RPC method changed from tasks/send to SendMessage, and the A2A-Version: 1.0 header is required
  • A PyPI mirror URL hardcoded into uv.lock is fixed by regenerating with UV_DEFAULT_INDEX
  • An A2A agent is a Starlette ASGI server, so it deploys to Cloud Run (a Docker container), not Cloud Functions (a single-function handler model)
  • On Cloud Run, the HOST_OVERRIDE environment variable sets the agent card’s public URL
  • The TypeScript A2A client needs the same v1.0 treatment: endpoint path, method name, role values, header
  • Because mocks and the real agent share the same client code and protocol, switching between test and production is a matter of changing a URL

References

Share this article