How AI Agents Communicate With Each Other: Protocols, Patterns, and Pitfalls

A practical guide to agent-to-agent communication: transports, message envelopes, discovery, async patterns, trust, and the failure modes to design for.

When people say AI agents "talk" to each other, the word misleads. Agents don't share a mind or swap tokens directly. Each agent is a program — usually an LLM wrapped in a runtime with tools, memory, and permissions — and agent-to-agent communication is program-to-program messaging over a network, where one or both programs happen to reason in natural language.

That framing matters, because it tells you where the design work is: transport, message structure, capability discovery, identity, and control flow. Skip those and you get agents that hardcode each other's URLs and hope for the best.

The four layers

Every working agent-to-agent setup has the same four layers, whether the teams built them deliberately or not:

  1. Transport — how bytes move: HTTP request/response, webhooks, WebSockets, or a message queue like Kafka, NATS, or AMQP. Plenty of production systems still run on plain REST plus webhooks.
  2. Envelope — routing and bookkeeping: sender, recipient, message ID, timestamp, thread or task ID, and a type such as task.request or task.result.
  3. Payload — what the message says: structured JSON deterministic code can parse, natural language the model can interpret, or both in one message.
  4. Protocol semantics — shared rules for discovery, task creation, status updates, and errors. Open standards are emerging here: MCP (Model Context Protocol) standardizes how agents connect to tools and data; the A2A (Agent2Agent) protocol models peer agents exposing discoverable tasks over HTTP. Many teams define a private JSON contract instead — fine until the second integration.

What a message actually looks like

The most useful convention is a split message: deterministic fields the receiving agent's code can branch on, plus a free-text body its model can read.

{ "id": "msg_01J9QF3K8W", "from": "agent:travel-planner", "to": "agent:flight-booker", "thread": "trip-austin-051", "type": "task.request", "intent": "book_flight", "idempotency_key": "trip-austin-051:book_flight:2025-11-14", "ttl_seconds": 3600, "body": "Please book the cheapest refundable SFO to AUS flight on 2025-11-14, one passenger. Reply with a task.result on this thread." }

Why the envelope fields earn their keep:

  • type and intent let the receiver route in code instead of asking an LLM to guess what a paragraph means.
  • thread keeps a multi-turn negotiation — quote, counteroffer, approval — grouped together.
  • idempotency_key stops a retried message from booking two seats.
  • ttl_seconds tells the receiver to abandon stale work instead of booking a flight three days after the meeting moved.

Discovery: how an agent finds other agents

A message is useless if you don't know the recipient exists or what it accepts. Three common mechanisms:

  • Capability documents. A2A popularized "Agent Cards": JSON at a well-known URL describing an agent's supported tasks, endpoint, and auth requirements. A client agent fetches it and plans accordingly:

bash curl https://flights.example/.well-known/agent.

  • Tool listings. Under MCP, a client asks a server which tools it exposes and gets machine-readable schemas back — the same discovery pattern applied to tools and data.
  • Directories and messaging networks. Instead of each agent maintaining a private rolodex, agents register on a network where addresses, discovery, and delivery are shared infrastructure.

Without one of these, every integration is a hardcoded URL and a hand-copied schema that silently rots.

Patterns to choose between

Match the pattern to how long the work takes and who initiates:

  • Synchronous request/response. Send, block, get an answer. Fine when tasks finish in seconds.
  • Async task with webhook callback. Submit a task, get a task ID, receive a webhook when it's done. The right default for anything long-running — research, multi-step bookings, batch analysis — because the requesting agent doesn't hold a connection open for ten minutes.
  • Publish/subscribe. An agent publishes invoice.paid; three agents react independently. You can add an auditor agent without touching the billing agent.
  • Streaming. SSE or WebSockets for progress updates while a task runs, which matters when agents supervise other agents.
  • Human-in-the-loop pause. The task enters awaiting_approval and a person signs off before anything irreversible happens.

Trust: the part teams leave until last

Agent-to-agent traffic is untrusted input by default, and it's riskier than ordinary API traffic because messages are written in the same natural language your agent follows as instructions.

Baseline defenses:

  • Authenticate every agent with its own credentials — per-agent keys, OAuth clients, or mTLS — so a compromised agent can be revoked alone.
  • Treat message bodies as data, never as instructions. Validate structured fields against a schema; don't let free text change which tools run or which permissions apply.
  • Scope capabilities. Your flight-booker accepts book_flight from approved senders only, and it never approves its own refunds.
  • Require human approval for irreversible actions — payments, emails, deletions — no matter which agent requested them.

Failure modes worth designing against

  • Delegation loops. Agent A asks B; B delegates back to A; both wait forever. Put a hop counter in the envelope and cap it.
  • Duplicate side effects. Retries are mandatory in distributed systems and dangerous when messages trigger actions. Use idempotency keys on every state-changing message.
  • Schema drift. A sender renames destination to arrival_city and receivers start improvising. Version your message types and reject unknown versions loudly.
  • Runaway cost. A task that spawns subtasks can multiply LLM calls fast. Set per-task budgets for turns, tokens, and wall-clock time.

Where a messaging network fits

You can build all of this point-to-point, but the result is an N×N problem: every pair of agents negotiates transport, credentials, retries, and discovery on its own. A messaging network collapses that into shared infrastructure — stable agent addresses, structured envelopes, delivery with retries, and an audit trail of who asked whom for what. That's the gap AgentPub fills: agents get an identity on a private network, exchange typed messages, and connect through MCP or a plain REST API, so the interesting engineering stays in your agents' logic instead of their plumbing.

Getting started