By Vaibhav Raizada, Kumar Kasimala, and Ashish Gite.
In our Engineering Energizers Q&A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Vaibhav Raizada, a senior software engineer at Salesforce, whose team faced a high-stakes AI engineering contradiction while building Assistive Authoring for Salesforce Prompt Builder: transforming open-ended instructions into complete, grounded prompt templates in under a minute without letting the non-deterministic LLM control routing, record identity, or the final output contract. One hallucinated field, corrupted ID, or malformed response could turn a plausible draft into a broken template.
Explore how Vaibhav’s team used Agent Graph, Agentforce’s orchestration framework, to build a deterministic LLM orchestration workflow across specialized nodes and four backend tools while preserving control over context, grounding, execution order, structured output, and human approval.
What mission did your team take on when AI prompt-template generation had to produce functional, grounded templates from natural-language requests?
The most dangerous failure was not an obviously bad response. It was a template that looked correct but broke when someone tried to use it. The LLM could invent a field, corrupt a record ID, return malformed metadata, or route the request down the wrong path. Any one of those failures could produce a convincing but unusable template.
The mission was to eliminate the expertise barrier around prompt-template creation without introducing those risks. Salesforce Prompt Builder lets admins create reusable prompts with dynamic placeholders populated at runtime using grounded CRM data. Building a good template, however, requires an admin to understand the org’s available data, use the correct merge fields and Jinja syntax, determine which information should become an input or data provider, structure the prompt effectively, and keep the result valid against the underlying schema.
For an experienced admin, that process can take at least 25 minutes. For someone new to prompt authoring, the blank canvas can become a wall. The goal was for an admin to describe a requirement in plain English and receive a functional template in under a minute, complete with valid inputs, data providers, merge fields, and metadata.
To achieve that 25x improvement, the team used Agent Graph to divide creation and refinement into focused nodes. This preserved the LLM’s ability to interpret language while keeping routing, identity, grounding, and structured output on deterministic rails.
What challenges emerged when orchestrating open-ended LLM requests across specialized Agent Graph nodes and four backend tools?
A single free-text box has to handle creation requests, refinements, vague instructions, and messages unrelated to authoring. Asking one large prompt to interpret every request, retain all supporting context, and control the entire workflow could overload the context window, weaken intent recognition, and send requests down the wrong path.
In response, the team created a router node that hands requests to three specialized sub-agents. generate_template creates a new template, refine_template revises existing content, and clarify_request handles non-authoring input without calling tools or producing template content. To carry the correct context into each path, Agent Graph uses state variables. External variables contain client-supplied information such as authoring mode, template type, existing content, template JSON, and inputs. Internal variables hold runtime information such as the next node, discovered data providers, and raw model output. Bound inputs pass only the required state into each tool, leaving the LLM to interpret the user’s request instead of reproducing structured payloads.

Router, sub-agent paths, and the shared tool pipeline behind Assistive Authoring.
With routing and context established, the graph coordinates four backend tools: fetch data providers, generate a template, refine a template, and transpose the result. A beforeReasoning hook fetches grounding before the LLM loop begins. Generation or refinement receives that context through state, and transpose becomes available only after template output exists. This sequence determines where the request goes, what context follows it, and which tool can run next. It also reduces latency by removing a model decision and tool round trip from the critical path. The result is a multi-step LLM orchestration workflow in which the graph controls the sequence and each tool performs one focused task.
What made deterministic LLM orchestration so difficult when routing, record identity, and structured output could not be left to the model?
The central danger was giving the LLM authority over decisions that could corrupt the workflow. If it chose the wrong route, altered an identifier, or improvised the response format, the system could produce an unusable template even when the language looked perfect. The team therefore established deterministic boundaries around routing, record identity, and structured output.
For routing, the model can signal intent through a state variable, but Agent Graph evaluates deterministic guards before changing nodes. The client also supplies the create-versus-refine mode. The LLM interprets fuzzy language, while the graph determines which transitions and tools are allowed. For record identity, a refined draft must map back to the existing template version, inputs, data providers, and parent records. Because LLMs can alter opaque identifiers, the model never controls them. It generates content using stable developer names, and a deterministic server-side transpose layer matches those elements to existing records and reapplies the correct IDs.
Finally, the editor needs structured LLM output rather than conversational text. The transpose step wraps the template metadata in sentinel tags, and the client extracts only the tagged content before parsing the JSON. If JSON parsing fails, the system returns an error instead of guessing.
The reusable pattern is to separate probabilistic work from structural control. The LLM interprets intent and generates language, while the graph governs routing, server-side code preserves identity, and the output contract determines what the client will accept.
What complexities did LLM grounding create when the system had to prevent hallucinated fields, inaccessible data, and invented references?
Grounding can span SObject fields, record snapshots, related lists, flows, Apex-backed providers, Einstein Search retrievers or any other invocable actions. Because each source can use different schemas, reference names, and merge-field conventions, sending everything directly to the LLM would inflate the context and increase the risk of invented or corrupted references.
To standardize that grounding context, the team built a discovery service that normalizes available providers into one consistent shape with the correct reference names attached. Discovery also runs in the current user’s context and filters fields based on field-level readability, surfacing only data the user can access and the prompt template can support. Data discovery is then separated from relevance selection. The available providers are reduced to an applicable set, which prompt-template generation treats as an allowlist. The model can select from validated data, but it cannot assume that a field exists or invent an object API name.
This creates a reusable LLM grounding sequence: discover the real data, normalize its representation, enforce the user’s permissions, reduce it to what is relevant, and constrain generation to that allowlist. Preventing LLM hallucinations required both deterministic discovery to establish trustworthy context and prompt engineering to enforce it as a hard boundary.
What safety challenges emerged when structured LLM output had to reach the editor without allowing the agent to save or execute anything?
The core human-in-the-loop safety constraint was that the output had to remain a proposed draft, never an action. The agent could generate or refine a template, but it could never save, publish, or execute one. Because the shared Copilot streaming adapter returns text rather than a structured object, the server wraps the template metadata in sentinel tags. The client extracts only the tagged payload and parses it into JSON. Malformed or incomplete metadata produces an error instead of being interpreted or partially applied.
The AI-generated content is then converted into typed editor blocks rather than injected as raw HTML. Conversational text uses framework interpolation that escapes the content, while pasted or incoming content is reduced to text. This prevents model output from becoming executable markup. Finally, the template appears as an accept-or-reject diff against the current content. Accepting it updates only the in-editor draft; saving remains a separate user action. The agent has no tool that can execute the template and cannot bypass the review gate.
Assistive Authoring can therefore generate a complete, grounded prompt template in under a minute without gaining authority over what becomes real. The model proposes, deterministic systems constrain and validate, and the admin decides whether the result is saved.
Learn more
- Stay connected — join our Talent Community!
- Check out our Technology and Product teams to learn how you can get involved.