The same question gets two different answers
So a single manual test proves nothing. You need a set of cases scored in aggregate, not a person trying it once and declaring it fine.
Agentforce testing and deployment is the release engineering for a system that is not deterministic: batch evaluations against expected topics and actions, guardrail and regression suites, a sandbox strategy that transfers to production, metadata deployment of agents and their actions, and staged rollout with a rollback path. You cannot assert an agent is correct. You can measure that it is good enough.
Trusted by 500+ organizations — including teams who needed a way to change a live agent without holding their breath.
Release Scope
Conventional Salesforce release practice assumes that the same input produces the same output. An agent breaks that assumption, and every habit built on it. The discipline has to change, not just the tooling.
So a single manual test proves nothing. You need a set of cases scored in aggregate, not a person trying it once and declaring it fine.
Tighten one topic description and a neighbouring intent starts routing elsewhere. Without a regression set, that lands in production undetected.
Evaluation runs are not free, in a sandbox or anywhere else. Test strategy has a cost line, which is why nobody should run the full suite fifty times a day.
Different data shape, different permissions, thinner content. Results that do not transfer are worse than no results, because they are believed.
Agents configured by hand in production cannot be diffed, reviewed or rolled back. The only record of what changed is somebody's memory.
Everyone tests that the agent does its job. Very few test that it declines the things it must decline — which is the failure with the worst consequences.
Agentforce testing works in layers, because an agent has deterministic and non-deterministic parts and they need different methods. The action bodies — Apex and Flow — are deterministic and take ordinary unit tests. The agent's behaviour — which topic it selects, which action it calls, how it populates inputs — is not, and is tested by running a batch of real utterances and scoring actual topic and action against expected. Salesforce provides the Testing Center for exactly this: batch evaluation against ground truth rather than one-off manual trials.
The mental shift is from correctness to acceptable behaviour at a measured rate. You cannot assert that an agent always answers correctly, in the way you can assert a Flow always sets a field. You can establish that it selects the right topic for ninety-something percent of a representative set, that it never breaches a guardrail in the adversarial suite, and that a proposed change moves those numbers in the right direction without breaking cases that passed before. That is a testable, shippable standard.
Salesforce provides the Testing Center with batch testing against expected topic and action, sandbox environments, and metadata deployment for agents and their components. Twopir Consulting provides the strategy and the plumbing: test sets built from real transcripts rather than imagination, evaluation criteria that mean something, guardrail and adversarial suites, a sandbox design whose results actually transfer, source control and deployment for agent metadata, and a staged rollout with a rollback somebody has rehearsed. You get the ability to change a live agent on a Tuesday without a war room.
One planning note we give every client: evaluation consumes the same AI capacity that production does, whether it runs in a sandbox or not. Test strategy therefore has a real budget, and the right design runs the fast deterministic layer often and the expensive evaluation layer at meaningful gates rather than continuously.
The cadence column matters as much as the rest. Running everything continuously is expensive; running nothing is worse; the design is in knowing which gate each layer belongs to.
| Layer | What it catches | When it runs |
|---|---|---|
| Apex & Flow unit tests | Wrong results, bulk failures, permission bypasses inside action bodies. | Every commit — fast, deterministic and cheap. |
| Batch evaluation | Wrong topic selected, wrong action called, inputs populated badly. | Before every promotion, and after any topic or instruction change. |
| Grounding & retrieval tests | The right topic answering from the wrong passage. | After any content change, and on a scheduled cycle. |
| Guardrail & adversarial suite | Refusals that do not hold, over-disclosure, authority exceeded. | Before every production release, without exception. |
| Regression suite | Previously working behaviour broken by an unrelated change. | Before every promotion — this is the one teams skip and regret. |
| Integration & latency tests | Slow or failing external calls inside a conversational turn. | Before release, and after any change to a connected system. |
| Production observation | Everything the test set did not anticipate. | Continuously, with new failure patterns fed back into the suite. |
Six deliverables. The point of all of them is that changing a live agent becomes routine rather than an event that needs three people watching a dashboard.
Built from real transcripts, not invented. Every case carries an expected topic and expected action, and the set deliberately includes the awkward phrasings, the typos and the two-questions-at-once.
Running the set through the Testing Center and turning the output into a decision: did this change improve topic and action selection, and by enough to justify promoting it?
Deliberately trying to make the agent misbehave — asking for other customers' data, requesting forbidden topics, talking it out of its own rules — and confirming that every stated refusal holds.
Environments whose results transfer. That means representative data shape, matching permissions and realistic content — plus a clear rule about what may only ever be verified in production.
Agent metadata, Apex, Flows, prompt templates and library configuration in source control and deployed like everything else — so a change can be reviewed, diffed, promoted and reverted.
First traffic is a slice, watched against the pre-release baseline, with a rollback path somebody has actually rehearsed rather than assumed — including what to do about conversations already in flight.
An agent release is rarely just the agent. These are the components a change usually spans, and what is special about promoting each one.
The agent definition and its topics move as metadata, which is what makes review and rollback possible. Teams configuring only in production have no diff and no way back — the first thing we change.
The capability behind the actions, deployed alongside the agent that calls them. Version drift between an action and its backing logic is a common cause of a change that works in sandbox and fails live.
Library configuration promotes, but the indexed content is environment-specific and has to be seeded and re-indexed per environment. Forgetting this is why sandbox retrieval results so often fail to transfer.
Deployed as metadata and re-verified as a release gate. A new action that needs access to a new object is a permission change, and it has to be reviewed rather than granted quietly to make a test pass.
Endpoints and credentials differ per environment, so the pipeline has to substitute them cleanly. Deciding what is stubbed in test and what is called live is a real design choice, not a detail.
Messaging and site configuration is where staged exposure is actually implemented. Releasing to one channel or one queue first is the cheapest risk control in the whole process.
Release verification depends on watching the right numbers against the pre-release baseline. If the observation window has no agreed criteria, “it looks fine” becomes the acceptance test.
The same repository and pipeline as the rest of your Salesforce work, with the fast deterministic tests wired in and the expensive evaluation layer triggered at gates rather than on every commit.
Three to six weeks to establish, then it is your team's to run. We deliberately build it to be operated by your people rather than by us.
Agree what good enough means, per topic: selection accuracy, guardrail compliance, latency ceilings. Without a written standard there is nothing for a gate to check against.
Functional cases from real transcripts, the adversarial guardrail suite, and the regression set — with expected topic and action recorded for every case.
Sandbox data shape, permissions and content brought close enough to production that results transfer — and an explicit list of what can only be verified live.
Agent metadata into source control, fast tests into CI, evaluation triggered at gates, and promotion made a repeatable operation rather than a careful afternoon.
Run a real change through the whole path including the rollback, so the first genuine incident is not the first time anybody has tried reverting.
Release engineering, integration testing and controlled go-lives are what our delivery practice has always been made of. Agents change the test method, not the discipline.
Working with Twopir Consulting was a game-changer for our organization. They designed and implemented a comprehensive Service Cloud + Experience Cloud solution that perfectly aligned with our customer support and partner portal needs. Their proactive approach, attention to detail, and post-go-live support have been outstanding.
Salesforce–MeetMax integration for a corporate networking organization.
Twopir helped us identify the right AI tools and select the best-fit platform for our needs. Their team provided end-to-end support — from consulting to implementation — delivering an AI-powered chatbot and automation system that improved our lead routing and customer engagement.
Automating email attachment processing and Salesforce data routing.
Most Agentforce work today is configured directly in production by people who would never do that to a Flow. The discipline exists already — it just has not been applied here yet.
Invented test cases test the phrasing you imagined. Real ones include the typos, the two-questions-at-once and the requests nobody planned for — which is where agents actually fail.
We try to break the refusals rather than confirming the happy path. A guardrail that has never been attacked is a hope, and it is the failure with the worst consequences.
Agent definitions, actions and templates in source control, deployed through your pipeline — so a change is reviewable, diffable and reversible like any other Salesforce work.
Evaluation consumes real capacity. Fast deterministic layers run constantly; expensive evaluation runs at meaningful gates. Ignoring that produces either a surprise bill or no testing at all.
A rollback plan nobody has executed is a document. We run one during the engagement, including what happens to conversations already in flight when you revert.
You stop testing for an exact output and start testing for correct behaviour at an acceptable rate. Batch evaluation runs a set of real utterances and scores whether the agent selected the expected topic and the expected action — those are checkable even when the wording varies. Guardrail cases are stricter and binary: a refusal either held or it did not. Response wording itself is assessed by sampling and human review rather than string matching, because exact-match assertions on generated text produce a suite that fails constantly and teaches everyone to ignore it.
Batch testing agents at scale. You supply a set of test cases — typically an utterance plus the expected topic and expected actions — and it runs them together and reports actual against expected, so you can evaluate many scenarios in one pass rather than typing questions by hand. It is the mechanism that makes the evaluation layer practical. What it does not supply is the test set itself, the judgement about what good looks like, or the discipline of running it before every promotion; those are the parts that determine whether it is useful.
Largely yes — agent definitions, topics, actions, Apex, Flows and prompt templates deploy as metadata through the same tooling as the rest of your Salesforce work, which is what makes source control and rollback possible. The parts that do not travel are environment-specific: indexed content has to be seeded and re-indexed per environment, credentials and endpoints differ, and channel deployments are usually configured per org. Getting those substitutions handled cleanly is most of what building the pipeline involves.
Big enough that a few cases flipping does not change the verdict, and small enough that you will actually run it. In practice that means covering every live topic with several phrasings each, including the boundary cases between adjacent topics, plus the full guardrail suite and everything that has broken in production before. Start smaller than feels rigorous and grow it from real failures — a set built entirely from imagination is both larger and less useful than one grown from what actually went wrong.
Yes. Running an agent consumes AI capacity whether it happens in a sandbox or in production, so evaluation runs draw on the same budget. That has a real design consequence: you cannot run a large batch evaluation on every commit the way you would run Apex tests. We structure the layers so the cheap deterministic tests run constantly and the expensive evaluation runs at meaningful gates — before a promotion, after a topic change, on a scheduled cycle — and we put a number on that cadence during planning rather than letting it surprise someone.
Carefully, and it is exactly why we rehearse it. Reverting the metadata is straightforward; the questions that need answering in advance are what happens to sessions already in progress, whether any action the new version took needs compensating, and whether reverting the agent also requires reverting a Flow or a permission change that shipped with it. Staged exposure limits the blast radius so the answer usually affects a handful of conversations rather than a day's traffic — which is the real argument for staging.
Tell us how your last agent change was tested and shipped. We will show you what the gaps cost, build the test sets and the pipeline around them, and rehearse a release with you — including the rollback.
Salesforce, CRM & AI delivery for growing and mid-market companies · Contact the team