Back to all notes
Artificial IntelligenceCompany note

A Design Question for Multi-Agent Systems: When Is Language the Wrong Interface?

Oyu Intelligence8 min read
A Design Question for Multi-Agent Systems: When Is Language the Wrong Interface?

An engineering essay on where free-form agent dialogue can lose intent, and when schemas, state machines, or typed messages may be easier to inspect.

A Design Question for Multi-Agent Systems: When Is Language the Wrong Interface?

AI Agent Communication

When several AI components coordinate a task, natural language is convenient but not always the easiest interface to validate. This article is an engineering argument, not a report on one definitive study: some handoffs may be clearer as schemas, explicit state, or typed events.

The Core Challenge: Structural Misalignment

One way to describe the issue is a mismatch between a model's internal numerical representations and the short text passed to another component. Text can omit constraints or collapse different states into similar wording, although that does not prove that a hidden numerical protocol is automatically safer or more useful.

For example, a message that says "the review is complete" may omit which record was reviewed, which checks ran, what evidence was used, and whether any exception remains. Explicit fields make those distinctions easier to validate.

The Translation Problem

Different states can produce similar natural-language summaries. The practical issue is not recovering a model's full internal representation; it is preserving the task state, constraints, evidence, and allowed next actions that another component needs.

Failure Modes to Test

These are practical failure modes to test in multi-step systems. They can arise from prompts, orchestration, state handling, tool contracts, or model behavior; free-form communication is one possible contributor rather than a proven single cause.

Intention Drift

One of the most problematic issues is what's called "semantic drift chain." As tasks pass from one agent to another, the original goal gradually warps with each handoff. This creates an accumulating error that compounds over time, similar to how a message changes in a game of telephone. The difference is that in AI systems, this drift can lead to agents pursuing objectives that diverge significantly from the intended goal.

The risk can be tested by carrying the original objective and acceptance criteria through each handoff, then comparing them with the final action.

Goal Misinterpretation

Ambiguous messages or incomplete state can contribute to several observable failures:

  • Agents getting stuck in recursive planning loops, continuously re-planning the same task without making progress
  • Redundant dialogue where agents repeat similar exchanges without advancing toward the goal
  • Creation of phantom subtasks that were never part of the original objective

The "Saying vs. Doing" Problem

An action-state gap occurs when a component reports completion without a corresponding tool result or state change. Completion should therefore be derived from verifiable state where possible, not from a generated status sentence alone.

This disconnect between reported status and actual execution can cascade through a multi-agent system, with subsequent agents making decisions based on false assumptions about what has been accomplished.

Role Drift

Multi-component systems may assign roles such as planner, executor, and validator. A prompt reminder alone does not guarantee that those boundaries persist through a long run; permissions and allowed operations can enforce the distinction more directly.

Test for role drift by checking which component proposed, executed, and approved each consequential action.

Choosing Explicit Interfaces

The interface should match the handoff. Natural language may suit an ambiguous request, while a schema, state machine, or typed event may suit a known operation.

Key Architectural Requirements

Four engineering patterns are worth evaluating:

1. Role Persistence

Bind permissions and allowed actions to the component or role instead of relying only on conversational memory.

2. Structured Communication

Use documented fields, identifiers, enums, and validation for information that has known structure. A human-readable summary can accompany the structured payload for inspection.

3. Inter-Agent State Synchronization

Persist shared goals, task state, evidence references, and version information in a durable store. Components should read the state they need and reject stale or conflicting updates.

4. Decoupled Modules

Separate proposal, execution, validation, and approval where the risk justifies it. This creates boundaries that can be tested and logged.

Implications for AI Development

These choices affect several parts of system design:

For System Design

Developers building multi-agent systems may need to reconsider fundamental architectural choices. Rather than defaulting to natural language as the communication medium, they should evaluate whether structured state exchange would better serve their use case.

For Performance and Reliability

Structured fields can make missing state and invalid transitions easier to detect. Reliability still depends on the implementation, model behavior, tools, tests, and operating controls.

For Scalability

As the number of handoffs grows, versioning, idempotency, conflict handling, and observability become more important. Scale should be demonstrated with load and failure tests rather than assumed from the protocol.

Practical Considerations

Structured communication also creates trade-offs:

Interpretability

Language is often easier for people to inspect. Schemas and event payloads should therefore have readable logs, labels, and links to the underlying evidence.

Hybrid Approaches

Rather than completely abandoning language, many systems may benefit from hybrid approaches—using structured communication for agent-to-agent interaction while maintaining language interfaces for human-agent communication.

Tooling and Infrastructure

Schemas, migrations, retries, tracing, and debugging tools add engineering work. That cost should be justified by the risk or repeatability of the handoff.

A Practical Decision Rule

Do not replace language merely because a system has several components. Start by identifying the information that another component must interpret without ambiguity. Then choose the smallest interface that preserves and validates it:

  • natural language for open-ended human requests and explanations;
  • schemas for known fields and constraints;
  • state machines for allowed transitions;
  • tool results for execution evidence; and
  • explicit approval records for consequential actions.

Conclusion

Intention drift, goal misinterpretation, action-state gaps, and role drift should first be treated as observable engineering failures. Better prompts may help, but explicit contracts, persisted state, validation, and tests are often easier to inspect.

Structured communication is not a universal replacement for language. It is another tool: useful when a handoff has known fields, allowed states, or safety requirements, while language remains useful for ambiguous input and human-facing explanation.

The useful design question is therefore specific: which information must be typed, validated, persisted, or approved, and which information benefits from flexible language?


For Oyu Intelligence projects, this is a working design principle rather than a product-performance claim: use structured handoffs where correctness and auditability matter, and keep language where interpretation is genuinely useful.

Oyu Intelligence

From the team

Oyu Intelligence

Editorial Team