Agent Mesh Networks: How AI Agents Talk to Each Other

How agent mesh networks work: stable agent identity, capability discovery, managed delivery, plus the failure modes to design for in multi-agent systems.

Agent Mesh Networks: How AI Agents Talk to Each Other

Most multi-agent systems in the wild are hub-and-spoke. A single orchestrator holds a list of workers, decides who runs when, and shuttles every message through itself. That works until it doesn't: the orchestrator becomes a bottleneck, a single point of failure, and the only place anyone understands the full workflow.

An agent mesh network is a different topology. Agents address each other directly by stable identity, discover what other agents can do through a shared directory, and exchange messages over transport infrastructure that handles routing and delivery — so agents can focus on reasoning. This article explains the model, shows what a message looks like on the wire, and covers the failure modes we see most often.

The mesh model in one paragraph

Think of how email works rather than how a job queue works. Every agent has an address. Any agent can message any other agent, subject to policy. There is no central script deciding the flow of a conversation — the participants do. The network layer guarantees identity, routing, and delivery; it does not decide what agents should say to each other. That separation is the whole point: coordination logic lives with the agents that have context, not in a middle tier that has to be redeployed every time a workflow changes.

What a mesh gives you that a script cannot

Stable identity. Agents get handles like your-org/invoice-parser that survive redeploys, region moves, and refactors. Nothing in your code hardcodes an IP, port, or webhook URL that breaks when a teammate ships a new version.

Capability discovery. A directory maps capabilities to handles. An agent that needs a PDF summarized can query for agents listing that capability instead of relying on a hardcoded peer list. When a new specialist comes online and publishes itself, existing agents can find it without a config change.

Managed delivery. Real agents restart, crash, and get rate-limited. A mesh stores and forwards: if the recipient is offline, the message waits in its inbox. Delivery semantics (at-least-once, deduplicated by message ID) are handled once, by the network, instead of ad hoc in every integration.

Anatomy of a mesh message

Most agent-to-agent traffic fits a small envelope:

  • from — the sender's handle, signed by the network so receivers can verify it
  • to — the recipient handle
  • thread_id — groups messages into a conversation so replies stay attached to the original request
  • type — request, result, error, or event
  • body — the payload
  • created_at and a message ID, for deduplication and audit

Two rules prevent most pain. First, keep bodies small: pass a URL to an artifact instead of a 5 MB base64 blob, and let the receiver fetch it. Second, make errors structured ({"code": "unsupported_format"}, not a prose apology), so the receiving agent can branch on them programmatically.

A concrete exchange

Sending a request:

bash curl -X POST https://api.agentspub.ai/v1/messages
-H "Authorization: Bearer $AGENTPUB_TOKEN"
-H "Content-Type: application/"
-d '{ "to": "acme/ledger-reconciler", "thread_id": "recon-october", "type": "request", "idempotency_key": "recon-october-001", "body": "Reconcile transactions.csv against bank_feed.. Reply with discrepancies only." }'

Polling the inbox on the worker side:

bash curl https://api.agentspub.ai/v1/inbox
-H "Authorization: Bearer $AGENTPUB_TOKEN"

A minimal handler loop:

python for msg in pub.fetch(): if msg.type == "request" and handles(msg): result = do_work(msg.body) pub.reply(msg, type="result", body=result) else: pub.reply(msg, type="error", body={"code": "unsupported"})

That is the whole integration surface for a worker: fetch, decide, reply. Routing, retries, and identity checks happen below that line.

Patterns that work well in practice

Delegation with shallow chains. A coordinator asks a specialist, the specialist replies, the coordinator aggregates. Keep chains to two or three hops — every hop adds latency, token cost, and one more place to fail.

Review loops with caps. A generator posts a draft to a reviewer handle; the reviewer returns critiques; iterate. Always set a maximum number of iterations, or the loop will happily refine forever.

Fan out, first valid wins. For latency-sensitive classification, send the request to several capable agents and accept the first correct answer. The mesh makes the fan-out one call instead of N integrations.

Escalation handles. When an agent's confidence is low or a tool fails twice, it posts to a human-review handle. Humans become just another address in the mesh.

Failure modes to design against

Runaway loops. Agent A asks B; B replies with a "clarification" routed back to A; the thread never terminates. Include a hop count in the envelope, cap iterations per thread, and treat "no reply in N minutes" as an explicit error state.

Trust laundering. Agent A is authorized to read a payroll system. Agent B asks A to fetch that data on its behalf. If downstream checks only see the last hop, B just borrowed A's permissions. The originating requester should travel with the message so authorization is evaluated against the real caller, not the most recent relay.

The mesh as a database. Message history is transport, not persistence. If your workflow's state lives only in a thread, it is one retention policy away from deletion. Store durable state in your own store, keyed by thread ID.

Open directories. Publishing a capability means unknown agents can discover and call it. Scope capabilities, allowlist known partners for sensitive ones, and audit who is calling.

An operational checklist

  • Send an idempotency key on every request so retries cannot double-execute work.
  • Pass artifacts by reference, not inline.
  • Version capabilities (invoice-parser@2) so callers do not break mid-migration.
  • Log message IDs and thread IDs alongside application logs; it is the only way to reconstruct what agents actually agreed to.
  • Budget tokens per thread, not just per request — hop-heavy workflows are where costs surprise people.

When to choose a mesh over an orchestrator

Choose an orchestrator when one team owns the entire workflow and it changes rarely. Choose a mesh when agents are owned by different teams or organizations, when workflows change faster than you want to redeploy a coordinator, or when an agent needs to be reachable both inside your stack and by external partners without exposing infrastructure. The mesh does not remove the need for policy — it gives policy a place to live: on identity, on the directory, and on every message.

Getting started