By Roopang Chauhan, Sundar Vedula, Peng-Wen Chen, and Nikhil Bojja.
In our Engineering Energizers Q&A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Roopang Chauhan, a software engineering architect on the Agentforce team, who helped architect Agentforce Memory, a capability designed to eliminate one of enterprise AI’s biggest limitations: stateless AI agents. Without persistent memory, enterprise AI agents force users to repeatedly explain who they are, what they have done, and what they need, creating unnecessary friction and longer conversations.
Explore how Roopang’s team tackled the engineering challenges behind Agentforce Memory, from establishing a common language for AI memory and evolving the architecture from proof of concept to pilot, to balancing semantic retrieval, context management, latency, developer ergonomics, and user trust while building a production-ready platform capability.
What is your team’s mission, and why did giving Agentforce persistent memory become such an important engineering challenge?
The mission is to build platform capabilities that make Agentforce agents increasingly intelligent, effective, and easier to work with. One capability quickly rose to the top of that list: Agentforce Memory, which brings persistent memory to enterprise AI agents.
The goal was to solve what the team calls the “Groundhog Day” problem of enterprise AI. Without memory, every interaction begins from a blank slate. Agents do not remember previous conversations, user preferences, or important context learned over time. Users must repeatedly explain the same information, answer the same follow-up questions, and rebuild context every time they return. As conversations become longer and more sophisticated, that repetition increases the number of conversational turns, slows task completion, and creates a frustrating experience.
Consider calling your bank. The best customer experiences do not start by asking you to explain everything from the beginning. They already know who you are, understand your recent activity, and can immediately focus on solving your problem. The goal was for Agentforce to deliver that same experience so users could spend less time rebuilding context and more time getting work done.
Customer feedback made clear that this capability was missing. Persistent memory had become a table-stakes feature for enterprise AI, and Agentforce needed it. That realization set the stage for a much bigger engineering challenge: building Agentforce Memory that customers could actually trust and use in production.
When your team first began building Agentforce Memory, what engineering challenge made it difficult for everyone to speak the same language about AI memory?
The first challenge was not implementation, rather, it was terminology. Across the AI industry, concepts such as semantic memory, episodic memory, procedural memory, short-term memory, and long-term memory are often defined differently depending on the vendor or research paper. That made productive architectural discussions surprisingly difficult because engineers frequently meant different things while using the same words.
Before building a scalable Agentforce Memory architecture, the team created a shared terminology document that defined each memory concept and its role within Agentforce. Once everyone understood the boundaries between memory itself and the subset of information ultimately injected into an agent’s runtime context, architectural discussions became much more productive. Establishing that common vocabulary became the foundation for every engineering decision that followed.
As Agentforce Memory evolved from proof of concept into pilot, what engineering challenges forced the architecture to evolve along the way?
The proof of concept showed that the idea worked, but productizing Agentforce Memory proved to be a much longer journey. Multiple platform components were involved, many owned by different teams, so building the complete architecture from day one would have delayed customer value. Instead, the team deliberately chose an incremental approach, delivering useful capabilities as soon as they were ready.
The process started by injecting conversation history and user profile information into an agent’s context because those capabilities already existed and could immediately improve response quality. As confidence grew, it became clear that recent information is not always the most relevant information. A user might mention a preference weeks earlier that becomes critical later, so the architecture evolved toward semantic retrieval, enabling Agentforce Memory to find relevant conversations rather than simply replaying the latest ones.
Initially these capabilities existed as explicit Agentforce actions. As adoption grew, the team recognized that customers should not have to manually configure every step. That led to moving Agentforce Memory directly into the platform so enabling it became as simple as adding context.memory block to agentscript, while the Reasoner automatically handled memory curation, retrieval, and context injection behind the scenes.
Why was not conversation history alone enough, and what engineering tradeoffs shaped how Agentforce Memory retrieves and injects context?
Simply injecting recent conversation history works initially, but it does not scale. Recent conversations are not always the conversations that matter most. If a user mentioned an important preference a month earlier, limiting retrieval to the latest messages would likely miss it entirely.
There was no single obvious solution. The team evaluated multiple approaches, including injecting all memories into the context, allowing the large language model to retrieve memories on demand, progressively revealing additional context, and performing semantic retrieval before reasoning begins. Each approach involved different tradeoffs around latency, context size, and retrieval quality.
Based on data science experiments and latency considerations, the team chose to perform keyword search before the reasoning process begins and inject memory only in case of a keyword match. That approach keeps the context window bounded while minimizing latency. Because retrieval happens during pre-orchestration, it executes alongside other pre-orchestration steps rather than delaying the user’s response. The result is a balance between richer agent context, efficient memory retrieval, and the responsiveness users expect from enterprise AI.

Why did building Agentforce Memory outside the Reasoner ultimately become the wrong architectural approach?
The original approach kept Agentforce Memory outside the Reasoner by exposing it through actions. While technically workable, it created unnecessary complexity for administrators because every sub-agent required additional actions, instructions, and configuration. The team eventually recognized that administrators were being asked to make far too many changes just to enable memory. Every additional action, instruction, and configuration increased both complexity and the perceived risk of breaking an existing agent.
Moving memory into the Reasoner resolved that problem. Enabling Agentforce Memory became a simple context.memory configuration instead of dozens of manual edits, which dramatically improved developer ergonomics while allowing memory retrieval to happen during pre-orchestration and minimizing latency at the same time. The result was a simpler experience for administrators and a better architecture for the platform.
As Agentforce Memory moved toward production, what engineering challenges made privacy, trust, and user control essential parts of the architecture?
As the team prepared Agentforce Memory for pilot, it became clear that building persistent memory was not only an engineering problem. It was also a trust problem. The initial assumption was that administrators would decide whether memory should be enabled for everyone. Legal and privacy requirements challenged that assumption, making clear that individual users also needed meaningful control over how their memories were created and used.
To address that, the team introduced user preferences that allow people to enable or disable Agentforce Memory globally or for individual agents, along with a Memory Management Console where users can review, edit, and manage their stored memories. Those capabilities are also being extended conversationally for channels where users do not interact through the standard interface.
Persistent memory is valuable not simply because an AI agent remembers more, but because it remembers the right information while remaining fast, transparent, and under the user’s control. Building Agentforce Memory meant treating privacy and trust as architectural requirements from the very beginning rather than capabilities added later.
Learn more
- Stay connected — join our Talent Community!
- Check out our Technology and Product teams to learn how you can get involved.