How AI Agents Communicate With Each Other: Protocols, Messages, and Failure Modes

A practical guide to how AI agents talk to each other: transports, MCP and A2A, message anatomy, async patterns, trust boundaries, and failure modes.

When two AI agents communicate, the interesting part is rarely the network call itself. Any HTTP client can move bytes between processes. The hard parts are everything around that call: how one agent describes what it wants, how the other gets discovered in the first place, how both sides track a conversation neither can fully "remember," and how each avoids treating the other's output as trusted instructions. This article breaks agent-to-agent communication into layers you can reason about and patterns you can apply directly.

Agents are not ordinary API callers

A conventional service call has a fixed contract: POST /refunds with a schema, executed deterministically. Agent communication differs in one fundamental way: the message is interpreted by a language model before anything is acted on. That buys flexibility — you can request things the sender's developer never hardcoded — but it removes guarantees. The receiving agent may misunderstand intent, invent a parameter, or obey an instruction hidden inside what was supposed to be data.

Everything below follows from that: make intent explicit, make context self-contained, and enforce trust boundaries in code rather than hoping the model behaves.

Three layers: transport, protocol, semantics

Transport is HTTP requests, webhooks, or queues. It determines whether calls are synchronous, whether the receiver must be online and publicly reachable, and who polls whom. Boring, and decisive.

Protocol covers envelopes, addressing, discovery, and task lifecycle. Your realistic options today are raw REST with JSON, MCP (built for connecting model hosts to tools and resources — most agent stacks use it for the "act on the world" half), and peer-agent protocols like A2A, which standardizes agent cards for capability discovery and task-oriented messaging between peers.

Semantics is what the payload means: structured fields a program can validate, natural language a model interprets, or both.

A practical production shape is a JSON envelope over HTTPS, with a structured payload for anything that triggers side effects and a free-text notes field for context the model should reason about.

What an agent-to-agent message should contain

A message that contains only prose forces the receiver to guess. A message that contains only fields starves the model of context. Use both:

{ "id": "msg_01J8ZQ4K", "thread_id": "thr_refund_8842", "from": "acme/triage", "to": "acme/refunds", "type": "request", "intent": "process_refund", "payload": { "order_id": "ORD-2291", "amount": 42.50, "currency": "USD", "reason": "damaged_item" }, "notes": "Customer sent photos of damage; support rep pre-approved within policy.", "reply_to": "acme/triage", "ttl_seconds": 3600, "idempotency_key": "refund:ORD-2291:2025-06-11" }

Why each part matters:

  • thread_id groups every message in one business situation. Agents have no shared memory; this is how both sides reconstruct the conversation.
  • intent states the goal in one unambiguous verb phrase, so the receiver's planner doesn't infer it from prose.
  • payload carries validated, machine-checkable data. Code on the receiving side should reject schema violations with a clear error rather than letting the model improvise.
  • notes carries nuance the schema can't. Treat it as untrusted text (more below).
  • ttl_seconds and idempotency_key prevent the two classic distributed-systems failures agents cause: acting on stale requests and retry storms.

Discovery: how agent A finds agent B

For two agents you own, hardcode the endpoints and move on. The problem gets interesting at three or more agents, or across organizations:

  • Static configuration. Fine for a fixed pipeline (triage → refunds → notification). Breaks silently when endpoints move.
  • Capability directories. Each agent publishes a card: name, endpoint, auth requirements, and the intents it handles. Senders ask "who can process a refund?" instead of naming hostnames.
  • Brokered delivery. A messaging layer holds the addressing relationship, so agents publish and subscribe by identity or topic while the network handles routing, retries, and access control.

The practical test: if renaming or relocating an agent requires editing every caller, you have an addressing problem, not a protocol problem.

Context has to travel inside the message

Agents don't share memory, a session store, or a vector database. Whatever the receiving agent needs must be in the message or reachable through references it can resolve. Two rules keep this sane:

  1. Summarize, don't forward transcripts. Send decisions, constraints, and open questions — not 40 turns of chat. Long transcripts burn tokens and bury the one constraint that matters.
  2. Pass artifacts by reference. Send a URL or ID for the order, ticket, or file, with enough metadata to authorize access.

And never put secrets in message content. Messages get logged, summarized, and echoed into model contexts on machines you don't control.

Async by default

Agents take seconds to minutes per step, so design for asynchrony:

bash curl -X POST "$AGENTPUB_API/v1/messages"
-H "Authorization: Bearer $AGENTPUB_TOKEN"
-H "Content-Type: application/"
-d '{ "to": "acme/refunds", "thread_id": "thr_refund_8842", "intent": "process_refund", "payload": { "order_id": "ORD-2291", "amount": 42.50 }, "idempotency_key": "refund:ORD-2291:2025-06-11" }'

The call returns immediately with a message ID; the refunds agent works on it and replies on the same thread_id. If the sender needs the answer to proceed, it waits on the thread (poll or webhook). If not, fire-and-forget with a TTL. Reserve synchronous request/response for short, read-only lookups.

Every peer message is untrusted input

This is the part teams underestimate. A peer agent's notes field is text that lands in your model's context, and text in context is instruction-shaped. If one agent in your network is compromised — or just buggy — its messages can steer yours.

Defenses that hold up in practice:

  • Validate payload against a schema before any side effect; ignore free-text instructions that conflict with the declared intent.
  • Scope credentials per peer: the refunds agent can call refund APIs, not read your CRM.
  • Require human or policy-engine approval for irreversible actions above a threshold.
  • Cap hops per thread so two agents can't negotiate forever.

Failure modes worth designing against

  • Loops. A asks B, B asks A. Cap with hop counts and TTLs.
  • Retry storms. Slow agent plus eager retries equals duplicate refunds. Idempotency keys on every mutating message.
  • Context truncation. A summarizer drops the constraint that mattered. Pin critical constraints as structured fields, not prose.
  • Intent drift. Over many turns, "process refund" becomes "cancel order." Restate the original intent in every message on the thread.

None of this requires exotic infrastructure. An envelope, a thread, a TTL, an idempotency key, and a trust boundary will carry you further than any framework choice.

Getting started