A practical guide to how AI agents talk to each other: transports, MCP and A2A, message anatomy, async patterns, trust boundaries, and failure modes.
When two AI agents communicate, the interesting part is rarely the network call itself. Any HTTP client can move bytes between processes. The hard parts are everything around that call: how one agent describes what it wants, how the other gets discovered in the first place, how both sides track a conversation neither can fully "remember," and how each avoids treating the other's output as trusted instructions. This article breaks agent-to-agent communication into layers you can reason about and patterns you can apply directly.
A conventional service call has a fixed contract: POST /refunds with a schema, executed deterministically. Agent communication differs in one fundamental way: the message is interpreted by a language model before anything is acted on. That buys flexibility — you can request things the sender's developer never hardcoded — but it removes guarantees. The receiving agent may misunderstand intent, invent a parameter, or obey an instruction hidden inside what was supposed to be data.
Everything below follows from that: make intent explicit, make context self-contained, and enforce trust boundaries in code rather than hoping the model behaves.
Transport is HTTP requests, webhooks, or queues. It determines whether calls are synchronous, whether the receiver must be online and publicly reachable, and who polls whom. Boring, and decisive.
Protocol covers envelopes, addressing, discovery, and task lifecycle. Your realistic options today are raw REST with JSON, MCP (built for connecting model hosts to tools and resources — most agent stacks use it for the "act on the world" half), and peer-agent protocols like A2A, which standardizes agent cards for capability discovery and task-oriented messaging between peers.
Semantics is what the payload means: structured fields a program can validate, natural language a model interprets, or both.
A practical production shape is a JSON envelope over HTTPS, with a structured payload for anything that triggers side effects and a free-text notes field for context the model should reason about.
A message that contains only prose forces the receiver to guess. A message that contains only fields starves the model of context. Use both:
{ "id": "msg_01J8ZQ4K", "thread_id": "thr_refund_8842", "from": "acme/triage", "to": "acme/refunds", "type": "request", "intent": "process_refund", "payload": { "order_id": "ORD-2291", "amount": 42.50, "currency": "USD", "reason": "damaged_item" }, "notes": "Customer sent photos of damage; support rep pre-approved within policy.", "reply_to": "acme/triage", "ttl_seconds": 3600, "idempotency_key": "refund:ORD-2291:2025-06-11" }
Why each part matters:
thread_id groups every message in one business situation. Agents have no shared memory; this is how both sides reconstruct the conversation.intent states the goal in one unambiguous verb phrase, so the receiver's planner doesn't infer it from prose.payload carries validated, machine-checkable data. Code on the receiving side should reject schema violations with a clear error rather than letting the model improvise.notes carries nuance the schema can't. Treat it as untrusted text (more below).ttl_seconds and idempotency_key prevent the two classic distributed-systems failures agents cause: acting on stale requests and retry storms.For two agents you own, hardcode the endpoints and move on. The problem gets interesting at three or more agents, or across organizations:
The practical test: if renaming or relocating an agent requires editing every caller, you have an addressing problem, not a protocol problem.
Agents don't share memory, a session store, or a vector database. Whatever the receiving agent needs must be in the message or reachable through references it can resolve. Two rules keep this sane:
And never put secrets in message content. Messages get logged, summarized, and echoed into model contexts on machines you don't control.
Agents take seconds to minutes per step, so design for asynchrony:
bash
curl -X POST "$AGENTPUB_API/v1/messages"
-H "Authorization: Bearer $AGENTPUB_TOKEN"
-H "Content-Type: application/"
-d '{
"to": "acme/refunds",
"thread_id": "thr_refund_8842",
"intent": "process_refund",
"payload": { "order_id": "ORD-2291", "amount": 42.50 },
"idempotency_key": "refund:ORD-2291:2025-06-11"
}'
The call returns immediately with a message ID; the refunds agent works on it and replies on the same thread_id. If the sender needs the answer to proceed, it waits on the thread (poll or webhook). If not, fire-and-forget with a TTL. Reserve synchronous request/response for short, read-only lookups.
This is the part teams underestimate. A peer agent's notes field is text that lands in your model's context, and text in context is instruction-shaped. If one agent in your network is compromised — or just buggy — its messages can steer yours.
Defenses that hold up in practice:
payload against a schema before any side effect; ignore free-text instructions that conflict with the declared intent.intent in every message on the thread.None of this requires exotic infrastructure. An envelope, a thread, a TTL, an idempotency key, and a trust boundary will carry you further than any framework choice.