By Sohini Arya and Manish Kumar Jha.
In our Engineering Energizers Q&A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Sohini Arya, AI Delivery Lead for Marketing Cloud, whose team pursued an audacious goal: build AI agents that could design, verify, test, and repair teams of other agents, even though the orchestrator was prohibited from writing a single agent file. The system had to transform an engineer’s request into a tested multi-agent team in approximately 15 to 30 minutes, down from two to four hours with the manual process.
Explore how the team reconciled conflicting architecture, failure-mode, and compliance decisions during multi-agent orchestration; taught verifier agents to distinguish warnings from genuine failures; and prevented self-healing AI agents from repeatedly repairing the wrong problem.
What mission drove your team to automate multi-agent AI team design and testing?
The manual workflow for building and testing multi-agent AI teams could erase the productivity gains engineers expected from using them. Engineers had to define YAML files, prompts, profiles, tools, models, budgets, and turn limits; test the assembled team; inspect each agent’s output; diagnose failures; modify the configuration; and run everything again. They also had to create agents capable of verifying the work of other agents, adding another layer they needed to design and validate themselves. That repetitive cycle trapped engineers in the mechanics of building agent teams instead of the problems those teams were meant to solve. Working with multiple uncoordinated agents increased the cognitive load because engineers still had to review, reconcile, and assemble their separate outputs.
In response, the team created Agent Designer to automate the work between the engineer’s initial request and final approval. Specialized agents analyze the architecture, failure modes, and compliance requirements before the orchestrator synthesizes their findings and pauses for review. Once approved, other agents write and verify the files, run the test scenario, diagnose failures, and perform bounded repair attempts. The orchestrator can dispatch and verify this work, but it cannot write an agent file itself.

From goal to governed agent outcome.
This governed workflow moves engineers from manually constructing every component to reviewing the architecture, approving the design, and evaluating the verification results. Within Marketing Cloud, it helped approximately 200 engineers begin adopting the manager-of-agents model.
Why is multi-agent AI governance harder than configuring one agent?
Adding agents created an authority and compatibility problem, not simply more configuration. Different models could join the same team, multiple members could attempt to modify the same artifacts, and a roster without one clearly defined coordinator could operate without dependable control. These risks emerged because a multi-agent AI team is a strict superset of a single-agent design. It must govern prompts, models, tools, profiles, and budgets across its full roster while defining coordination, compatibility, authority, and responsibility boundaries.
To address those dependencies, Agent Designer contains an orchestrator and seven specialists, each with its own model, budget, and turn limit. Teams created through Claude Unleashed must also define exactly one coordinator and pass cu teams validate to confirm backend and model compatibility. Compatibility was only one part of AI agent governance. The agent specification did not include a native write-scope field, so the team enforced each member’s write boundaries within its prompt. This created a deterministic boundary around decisions that could not be left to the model, preventing agents from modifying overlapping artifacts or assuming authority they had not been assigned.
Why does parallel multi-agent orchestration create hidden design conflicts?
Three specialists could produce independently valid answers that combined into an unbuildable design. Because the architecture, failure-mode, and compliance analysts worked without seeing one another’s findings, their contradictions appeared only when the work converged.
For example, the architecture analyst might select a topology that the compliance analyst later determined would exceed its budget. The failure-mode analyst could identify a reliability requirement that changed how responsibilities needed to be divided across the team. Since each analysis ran independently, the agents could not resolve those conflicts while working.
The orchestrator therefore had to detect and reconcile their contradictions before presenting one viable architecture. Rather than restarting the complete analysis, it uses bounded remediation to incorporate the required correction into the synthesis. The engineer then reviews the roles, tools, budgets, and AI agent verification requirements before authorizing any file creation.

Escalating high-impact decisions while automating routine work.
How did your team enforce budgets across different multi-agent architectures?
A budget that looked safe for one agent could multiply across every item, step, and team member until the workflow lacked the resources needed to finish testing or repair. One static spending limit could not accurately govern every agent architecture. To address those differences, Agent Designer stores budgets, turn limits, and cost ceilings for every tier and construct type in one configuration file. It then applies the formula required by the topology: a flat check for a single agent, items × cost × steps for a swarm, and whole-roster ceilings for a team.
Those calculations also had to fit within the orchestrator’s own resource limits. Its 500-turn budget is divided across analysis, human approval, file generation, testing, and self-healing retries. Consuming too much during an early phase could leave a partially built team without enough capacity for verification. Centralizing these limits removed the need for engineers to remember every safeguard. Topology-specific calculations also prevented a design from committing to a resource plan it could not complete, making budget enforcement part of the multi-agent architecture instead of another manual check.
How did your team prevent AI agent verification from approving the wrong designs?
Moving verification from engineers to agents meant the auditor itself could be wrong. A permissive verifier could approve a genuine violation, while an overzealous one could bury a sound design beneath false failures and prevent it from reaching testing.
To make AI agent verification machine-readable, the team required each verifier to return structured PASS, WARN, or FAIL blocks instead of free-form explanations. The orchestrator could then parse and compare the findings before deciding whether the design could continue, needed remediation, or required human attention. However, a structured format could not guarantee that the underlying judgment was correct. Rules such as “warn rather than fail when uncertain” and “do not fail a Tier 4 design solely for using bypassPermissions” help distinguish reviewable concerns from genuine violations.
Because no verifier could be trusted in isolation, Agent Designer also performs a meta-check across their outputs. Each verifier must satisfy its own contract, while the orchestrator determines whether the agents agree and whether their combined conclusions match the proposed design.
How did your team stop self-healing AI agents from repairing the wrong problem?
An autonomous repair loop could consume turns and money while making a failed system worse. If it misdiagnosed the root cause, each fix could introduce another change and move the team farther from the expected behavior without producing an obvious error.

A dependable agent continuously plans, acts, evaluates, and improves.
To identify what had failed, Agent Designer’s test doctor performs test-failure classification, categorizing problems as violations, permission blocks, tool errors, boot errors, soft failures, infinite loops, or state corruption. Most categories provide recognizable signals, but soft failures can execute normally and still return the wrong result. Detecting them requires comparing the output point by point with the expected behavior. Because diagnosis was not always deterministic, stopping behavior became as important as repair behavior. The team limited the self-healing process to two fix attempts and a $3 spending ceiling. These bounded retry and recovery controls ensure that Agent Designer stops and escalates if the team still fails instead of continuing to pursue a potentially incorrect diagnosis.
Agent Designer does not remove engineers from multi-agent AI development. It elevates them to the control points that matter: defining the outcome, approving the architecture, reviewing verification results, and deciding when autonomous repair must stop.
Learn More
- Stay connected by joining our Talent Community.
- Explore our Technology and Product teams to see how you can get involved.