A Design Question for Multi-Agent Systems: When Is Language the Wrong Interface?

An engineering essay on where free-form agent dialogue can lose intent, and when schemas, state machines, or typed messages may be easier to inspect.
A Design Question for Multi-Agent Systems: When Is Language the Wrong Interface?

When several AI components coordinate a task, natural language is convenient but not always the easiest interface to validate. This article is an engineering argument, not a report on one definitive study: some handoffs may be clearer as schemas, explicit state, or typed events.
The Core Challenge: Structural Misalignment
One way to describe the issue is a mismatch between a model's internal numerical representations and the short text passed to another component. Text can omit constraints or collapse different states into similar wording, although that does not prove that a hidden numerical protocol is automatically safer or more useful.
For example, a message that says "the review is complete" may omit which record was reviewed, which checks ran, what evidence was used, and whether any exception remains. Explicit fields make those distinctions easier to validate.
The Translation Problem
Different states can produce similar natural-language summaries. The practical issue is not recovering a model's full internal representation; it is preserving the task state, constraints, evidence, and allowed next actions that another component needs.
Failure Modes to Test
These are practical failure modes to test in multi-step systems. They can arise from prompts, orchestration, state handling, tool contracts, or model behavior; free-form communication is one possible contributor rather than a proven single cause.
Intention Drift
One of the most problematic issues is what's called "semantic drift chain." As tasks pass from one agent to another, the original goal gradually warps with each handoff. This creates an accumulating error that compounds over time, similar to how a message changes in a game of telephone. The difference is that in AI systems, this drift can lead to agents pursuing objectives that diverge significantly from the intended goal.
The risk can be tested by carrying the original objective and acceptance criteria through each handoff, then comparing them with the final action.
Goal Misinterpretation
Ambiguous messages or incomplete state can contribute to several observable failures:
- Agents getting stuck in recursive planning loops, continuously re-planning the same task without making progress
- Redundant dialogue where agents repeat similar exchanges without advancing toward the goal
- Creation of phantom subtasks that were never part of the original objective
The "Saying vs. Doing" Problem
An action-state gap occurs when a component reports completion without a corresponding tool result or state change. Completion should therefore be derived from verifiable state where possible, not from a generated status sentence alone.
This disconnect between reported status and actual execution can cascade through a multi-agent system, with subsequent agents making decisions based on false assumptions about what has been accomplished.
Role Drift
Multi-component systems may assign roles such as planner, executor, and validator. A prompt reminder alone does not guarantee that those boundaries persist through a long run; permissions and allowed operations can enforce the distinction more directly.
Test for role drift by checking which component proposed, executed, and approved each consequential action.
Choosing Explicit Interfaces
The interface should match the handoff. Natural language may suit an ambiguous request, while a schema, state machine, or typed event may suit a known operation.
Key Architectural Requirements
Four engineering patterns are worth evaluating:
1. Role Persistence
Bind permissions and allowed actions to the component or role instead of relying only on conversational memory.
2. Structured Communication
Use documented fields, identifiers, enums, and validation for information that has known structure. A human-readable summary can accompany the structured payload for inspection.
3. Inter-Agent State Synchronization
Persist shared goals, task state, evidence references, and version information in a durable store. Components should read the state they need and reject stale or conflicting updates.
4. Decoupled Modules
Separate proposal, execution, validation, and approval where the risk justifies it. This creates boundaries that can be tested and logged.
Implications for AI Development
These choices affect several parts of system design:
For System Design
Developers building multi-agent systems may need to reconsider fundamental architectural choices. Rather than defaulting to natural language as the communication medium, they should evaluate whether structured state exchange would better serve their use case.
For Performance and Reliability
Structured fields can make missing state and invalid transitions easier to detect. Reliability still depends on the implementation, model behavior, tools, tests, and operating controls.
For Scalability
As the number of handoffs grows, versioning, idempotency, conflict handling, and observability become more important. Scale should be demonstrated with load and failure tests rather than assumed from the protocol.
Practical Considerations
Structured communication also creates trade-offs:
Interpretability
Language is often easier for people to inspect. Schemas and event payloads should therefore have readable logs, labels, and links to the underlying evidence.
Hybrid Approaches
Rather than completely abandoning language, many systems may benefit from hybrid approaches—using structured communication for agent-to-agent interaction while maintaining language interfaces for human-agent communication.
Tooling and Infrastructure
Schemas, migrations, retries, tracing, and debugging tools add engineering work. That cost should be justified by the risk or repeatability of the handoff.
A Practical Decision Rule
Do not replace language merely because a system has several components. Start by identifying the information that another component must interpret without ambiguity. Then choose the smallest interface that preserves and validates it:
- natural language for open-ended human requests and explanations;
- schemas for known fields and constraints;
- state machines for allowed transitions;
- tool results for execution evidence; and
- explicit approval records for consequential actions.
Conclusion
Intention drift, goal misinterpretation, action-state gaps, and role drift should first be treated as observable engineering failures. Better prompts may help, but explicit contracts, persisted state, validation, and tests are often easier to inspect.
Structured communication is not a universal replacement for language. It is another tool: useful when a handoff has known fields, allowed states, or safety requirements, while language remains useful for ambiguous input and human-facing explanation.
The useful design question is therefore specific: which information must be typed, validated, persisted, or approved, and which information benefits from flexible language?
For Oyu Intelligence projects, this is a working design principle rather than a product-performance claim: use structured handoffs where correctness and auditability matter, and keep language where interpretation is genuinely useful.

Read Next
Related notes

Adversarial Poetry and AI Safety: What a 25-Model Study Found
In a preprint, 20 hand-written poetic prompts produced a 62% average attack-success rate across 25 models; a larger 1,200-prompt conversion test averaged about 43%.

PROMPTFLUX: What Google’s Report Says About AI-Assisted Malware
Google reported an experimental VBScript dropper that queried Gemini to rewrite and obfuscate its code. The sample was still in development, and Google did not report a successful device or network compromise.

Automation, AI Workflows, and Agents: A Practical Boundary
A concise way to choose between fixed rules, model-assisted steps, and systems that select among tools—without calling every workflow an agent.