Containment is flat and nobody knows which slice is stuck
A single blended number hides everything. Containment by topic almost always shows two or three areas dragging the whole figure down.
Agentforce optimization is the measured work of improving an agent that is already in production — raising containment and answer accuracy, cutting unnecessary escalations and latency, and controlling what each conversation consumes. It runs on evidence from real sessions, not opinions from a demo. Every change is scored before and after.
Trusted by 500+ organizations — including teams whose agent went live, plateaued, and needed someone to find out why.
Optimization Scope
A plateau is usually four or five separate causes stacked on top of each other, and they cannot be untangled without evidence. Guessing is what makes a plateau permanent.
A single blended number hides everything. Containment by topic almost always shows two or three areas dragging the whole figure down.
Each handoff looks individually reasonable. In aggregate the agent is escalating a category it could handle, because one instruction is too cautious.
Latency is usually one chatty action or a slow external call. Nobody notices until abandonment shows up next to average response time.
Conversations that take nine turns and four actions instead of three and one cost multiples more. Design decides the bill more than volume does.
Correct answers delivered in four paragraphs, or three clarifying questions before a simple reply. CSAT drifts down while accuracy looks fine.
Products changed, articles were edited, a Flow was updated. The agent was never wrong at launch — it is wrong now, for reasons nobody logged.
Agentforce optimization is a measured improvement cycle on a live agent. You observe what production is doing, diagnose which of a small number of causes is responsible, change one thing, score it against a fixed test set, and keep the change only if the score improved without breaking something else. It is closer to performance engineering than to configuration — the discipline is in the measurement, not in the cleverness of the change.
Every improvement in an Agentforce agent comes from one of four places, and knowing which is the whole job: the topic and instruction layer (it chose the wrong job or behaved wrongly inside the right one), action quality (it chose correctly but the action failed, was slow, or returned something unusable), grounding quality (it retrieved the wrong passage, or none), and data quality (the record it read was incomplete or duplicated). Teams that skip the diagnosis rewrite instructions for a grounding problem and conclude that AI does not work. Where the agent has never worked well, a readiness assessment is often the better starting point than a tuning cycle.
Salesforce provides the observability surface — session data, topic distribution, escalation and latency reporting, and the Testing Center for batch evaluation. Twopir Consulting provides the method: a fixed test set so change is measurable, the diagnostic discipline to attribute a symptom to the right layer, the fix itself across configuration, grounding, actions or data, and a review cadence that keeps it from drifting back. You get containment and accuracy that move, escalations that fall for the right reasons, and a cost per resolved conversation you can defend.
One caution we give every client: containment on its own is a misleading target. An agent can contain a conversation and leave the work undone, or contain it by wearing the customer out. We measure containment alongside accuracy, completion and CSAT, because optimising a single number is how an agent gets worse while its dashboard improves.
The right-hand column is where the work happens. Most wasted optimisation effort comes from attributing a symptom to the wrong layer and tuning something that was never the cause.
| What you see | What it usually means | Where the fix lives |
|---|---|---|
| Low containment in one topic | That topic's scope is too narrow, or its action fails and the agent falls back to a human. | Topic scope and action reliability. |
| Containment high, CSAT low | The agent is holding conversations it resolves badly — customers give up rather than escalate. | Answer quality and escalation triggers. |
| Inconsistent answers to similar questions | Two topics both plausibly match, so classification flips with phrasing. | Topic boundaries and classification descriptions. |
| Confidently wrong answers | Retrieval is returning a real but stale or contradictory passage. | Grounding and content. |
| High latency on some turns | One slow action — usually an external call without a timeout, or too many actions per turn. | Action design and integration pattern. |
| Consumption above forecast | Conversations run long, or each turn triggers more actions than it needs. | Instruction brevity, action count, and turn design. |
| Performance decaying over months | Content, products or automation changed underneath an agent nobody re-tested. | Review cadence and ongoing ownership. |
Each of these is a separate hypothesis with a separate test. We change one at a time because a batch of five changes tells you nothing about which one worked.
The opening engagement. We segment every metric by topic and channel, read a statistically useful sample of real sessions, and produce a ranked list of causes with the expected gain from each.
Where classification is unstable or behaviour is off-policy. Boundaries redrawn against real phrasing, instructions rewritten as testable rules, and each change scored on its own.
When the agent chose right and still answered badly. We test what retrieval returns per question and check whether actions succeed, time out or return something the agent cannot use.
Getting the moment right in both directions. Too eager destroys containment; too stubborn destroys satisfaction. Both are tuned against outcomes, not against a target number.
Cost and speed are mostly design outcomes. Fewer turns, fewer actions per turn, tighter responses and faster integrations all move both at once — without reducing what the agent can do.
The dashboards that make the next cycle possible, plus the meeting that uses them. Optimisation without a cadence is a one-off project that decays the month after it ends.
Optimisation is only as good as its instrumentation. These are the sources we connect, and what each one is uniquely able to tell you.
The only source that shows why something went wrong rather than that it did. Reading two hundred real sessions produces a better backlog than any dashboard, and it is where the diagnostic always starts.
Topic distribution, escalation rates and latency at the aggregate level. It tells you where to look and how big a problem is; it cannot tell you what caused it. Both halves matter.
Batch runs of a fixed test set, scoring expected topic and action against what actually happened. This is what makes “the change helped” a measurement rather than an impression — and it consumes budget, so runs are planned.
Whether the case actually closed, whether it reopened a week later, whether the lead converted. Conversation-level metrics flatter an agent; record-level outcomes are where the truth is.
Usage per conversation and per topic, so cost can be attributed to design decisions rather than absorbed as a fixed platform expense. Without it, nobody can justify making the agent less chatty.
The people receiving escalations know which handoffs were premature and which arrived too late. A lightweight feedback loop from them is the cheapest signal in the whole programme, and the most ignored.
The counterweight to containment. Tracked per topic rather than blended, so a well-contained but frustrating area cannot hide behind a healthy average across everything else.
What changed, when, and what it was expected to do. Without it, a metric that moves three weeks after four edits is unattributable — and the team relearns the same lesson every quarter.
A first diagnostic runs two to three weeks. After that, tuning cycles are short and repeating — typically two to four weeks each, with a review at the end of every one.
Segment the metrics that are currently blended, build the fixed test set from real sessions, and record where the agent stands today across every measure we intend to move.
Read sessions and attribute each failure to topic, action, grounding or data. The output is a ranked backlog where every item names its cause, not just its symptom.
Take the highest-expected-gain item, state what should improve and by roughly how much, and change only that — in a sandbox, with the current behaviour preserved.
Run the same test set. Confirm the target measure moved and that previously-passing cases still pass — then promote, or discard the change and record why.
Report what moved, log what changed, and start the next cycle from the updated backlog — with the test set growing as new failure patterns appear in production.
Diagnosing a live system, attributing causes and improving it against a baseline is the same discipline whether the system is an agent, a scoring model or a CRM.
The AI readiness engagement gave us a clear roadmap to operationalize AI across our processes. The team built intelligent ‘Next Best Action’ capabilities using scoring, engagement, and fit models, which significantly improved how we prioritize and interact with prospects. Their understanding of both CRM and AI-driven decisioning made a real difference in aligning our systems with business outcomes.
Automating email attachment processing and Salesforce data routing.
We engaged Twopir Consulting to conduct a Salesforce audit, and their structured, insight-driven approach exceeded our expectations. Their team quickly understood our complex processes, identified critical gaps, and provided clear, actionable recommendations. The audit improved our data accuracy and streamlined workflows.
Salesforce–MeetMax integration for a corporate networking organization.
Optimisation is the easiest service to deliver badly, because a confident-sounding change and a real improvement look identical without a scoreboard.
Built from real sessions and held stable across cycles. Without it, every claim about improvement is an opinion, and the agent's history becomes unknowable.
Topic, action, grounding or data. Naming the layer prevents the most common waste in this work: rewriting instructions for a problem that lives in the knowledge base.
Slower in the short run and far faster over a quarter, because you learn which changes actually work instead of accumulating edits nobody can attribute.
If the answer is a content rewrite, a Flow change or a data resolution rather than an instruction, we do that work — because an optimisation practice that only touches prompts hits a ceiling fast.
We track consumption per resolved conversation and treat it as something the architecture controls, so the economics improve alongside the quality rather than against it.
There is no industry number worth quoting you, because containment depends almost entirely on what mix of requests reaches the agent. An agent scoped to documented, transactional questions will contain a high share of them; the same agent exposed to everything, including requests it was never designed for, will look far worse on an identical build. That is why we measure containment per topic against your own baseline rather than against a benchmark — the only honest comparison is your agent last month.
Usually not. In our experience the majority of wrong answers trace to grounding — the agent retrieved a real passage that is stale, contradicted by another article, or written so ambiguously that it supports the wrong reading. The diagnostic step is simple: look at what was retrieved for the failing question. If the passage is wrong, no instruction rewrite will save it. If the passage was right and the answer was still wrong, then you are in instruction or topic territory.
Design, mostly. Whatever commercial model you are on, consumption follows how many turns a conversation takes and how many actions each turn triggers — so an agent that asks three clarifying questions before answering costs several times one that reads context it already has. The levers are shorter instructions, better input mapping so nothing is re-asked, fewer redundant actions per turn, and resolving conversations in fewer exchanges. Also budget for testing: evaluation runs consume capacity, so a heavy tuning programme has a real cost of its own.
The diagnostic takes two to three weeks and usually surfaces one or two changes that can ship immediately — a topic boundary, an over-cautious instruction, a badly mapped action input. Those show up in the numbers within days of promotion. Changes that depend on content remediation or data resolution take longer, because the fix is not in the agent at all. What we will not do is ship a batch of changes in week one and claim credit for whatever moves; attribution requires patience.
Because something around it changed. Agents depend on knowledge articles, records, Flows, integrations and product facts that all move independently — an article edited by a content owner, a new SKU with no documentation, a Flow updated by an admin, a policy changed by legal. The agent was correct at launch and is now quoting a world that no longer exists. This is why we treat a re-test cadence as part of the service rather than a nice-to-have: drift is the default state, not an incident.
Yes, and most should eventually. The method is learnable: keep a fixed test set, diagnose to a layer before changing anything, change one variable at a time, re-score, and log what you did. What a first engagement usually supplies is the diagnostic instinct — recognising that inconsistent answers mean topic overlap, that confident wrongness means grounding, that a latency spike means one slow action. We build the test set and the dashboards to be handed over, and several clients run their own cycles after two rounds with us.
A performance diagnostic works from evidence you already have. We segment your metrics, read real sessions, attribute each failure to a layer, and hand you a ranked backlog with the expected gain from each item.
Salesforce, CRM & AI delivery for growing and mid-market companies · Contact the team