Agentforce · Optimization & Performance

Live is the starting line, not the finish.

Agentforce optimization is the measured work of improving an agent that is already in production — raising containment and answer accuracy, cutting unnecessary escalations and latency, and controlling what each conversation consumes. It runs on evidence from real sessions, not opinions from a demo. Every change is scored before and after.

The Optimization Loop
PRODUCTION SIGNALS Sessions Transcripts · Outcomes · Drop-offs Measures Containment · Accuracy · CSAT Topic Mix Escalations Latency & Spend TWOPIR TUNING LAYER Diagnose Topic · Grounding Action · Data Change One Thing Isolated · Reversible Hypothesis first Re-Score Fixed test set Regression check NO CHANGE SHIPS WITHOUT A BEFORE AND AN AFTER 2πr WHAT MOVES More Contained Resolved without a human, correctly More Accurate Right topic, right source, right answer Cheaper Per Job Fewer wasted turns and actions MEASURE · DIAGNOSE · CHANGE · RE-SCORE · REPEAT
12+
Years CRM & AI delivery
250+
Deployments delivered
500+
Clients served
40+
Consultants & engineers

Trusted by 500+ organizations — including teams whose agent went live, plateaued, and needed someone to find out why.

Amberscript
Kacific
Spinify
Zarraffa’s Coffee
Ultra Consultants
RChilli

Optimization Scope

  • Salesforce Partner
  • Containment
  • Answer Accuracy
  • Escalation Design
  • Latency
  • Consumption Control
  • Observability
  • CSAT
Where Agents Plateau

It worked at launch. Then the numbers stopped moving

A plateau is usually four or five separate causes stacked on top of each other, and they cannot be untangled without evidence. Guessing is what makes a plateau permanent.

Containment is flat and nobody knows which slice is stuck

A single blended number hides everything. Containment by topic almost always shows two or three areas dragging the whole figure down.

Escalations are high but not obviously wrong

Each handoff looks individually reasonable. In aggregate the agent is escalating a category it could handle, because one instruction is too cautious.

Answers are slow enough that people leave

Latency is usually one chatty action or a slow external call. Nobody notices until abandonment shows up next to average response time.

Consumption is higher than the business case assumed

Conversations that take nine turns and four actions instead of three and one cost multiples more. Design decides the bill more than volume does.

Customers are technically satisfied and quietly annoyed

Correct answers delivered in four paragraphs, or three clarifying questions before a simple reply. CSAT drifts down while accuracy looks fine.

It has drifted since launch and nobody noticed

Products changed, articles were edited, a Flow was updated. The agent was never wrong at launch — it is wrong now, for reasons nobody logged.

What It Means

Optimization is measurement with a method

Agentforce optimization is a measured improvement cycle on a live agent. You observe what production is doing, diagnose which of a small number of causes is responsible, change one thing, score it against a fixed test set, and keep the change only if the score improved without breaking something else. It is closer to performance engineering than to configuration — the discipline is in the measurement, not in the cleverness of the change.

Every improvement in an Agentforce agent comes from one of four places, and knowing which is the whole job: the topic and instruction layer (it chose the wrong job or behaved wrongly inside the right one), action quality (it chose correctly but the action failed, was slow, or returned something unusable), grounding quality (it retrieved the wrong passage, or none), and data quality (the record it read was incomplete or duplicated). Teams that skip the diagnosis rewrite instructions for a grounding problem and conclude that AI does not work. Where the agent has never worked well, a readiness assessment is often the better starting point than a tuning cycle.

Salesforce provides the observability surface — session data, topic distribution, escalation and latency reporting, and the Testing Center for batch evaluation. Twopir Consulting provides the method: a fixed test set so change is measurable, the diagnostic discipline to attribute a symptom to the right layer, the fix itself across configuration, grounding, actions or data, and a review cadence that keeps it from drifting back. You get containment and accuracy that move, escalations that fall for the right reasons, and a cost per resolved conversation you can defend.

One caution we give every client: containment on its own is a misleading target. An agent can contain a conversation and leave the work undone, or contain it by wearing the customer out. We measure containment alongside accuracy, completion and CSAT, because optimising a single number is how an agent gets worse while its dashboard improves.

Metric to Fix

What each number is actually telling you

The right-hand column is where the work happens. Most wasted optimisation effort comes from attributing a symptom to the wrong layer and tuning something that was never the cause.

Agentforce performance metrics, what each usually indicates, and which layer to change
What you seeWhat it usually meansWhere the fix lives
Low containment in one topicThat topic's scope is too narrow, or its action fails and the agent falls back to a human.Topic scope and action reliability.
Containment high, CSAT lowThe agent is holding conversations it resolves badly — customers give up rather than escalate.Answer quality and escalation triggers.
Inconsistent answers to similar questionsTwo topics both plausibly match, so classification flips with phrasing.Topic boundaries and classification descriptions.
Confidently wrong answersRetrieval is returning a real but stale or contradictory passage.Grounding and content.
High latency on some turnsOne slow action — usually an external call without a timeout, or too many actions per turn.Action design and integration pattern.
Consumption above forecastConversations run long, or each turn triggers more actions than it needs.Instruction brevity, action count, and turn design.
Performance decaying over monthsContent, products or automation changed underneath an agent nobody re-tested.Review cadence and ongoing ownership.
Diagnose to a layer before changing anything. A grounding problem never gets fixed by rewriting instructions.
What We Tune

Six levers, pulled against a scoreboard

Each of these is a separate hypothesis with a separate test. We change one at a time because a batch of five changes tells you nothing about which one worked.

Performance Diagnostic

The opening engagement. We segment every metric by topic and channel, read a statistically useful sample of real sessions, and produce a ranked list of causes with the expected gain from each.

  • Metrics segmented by topic, channel and intent
  • Session sampling and failure classification
  • Attribution to topic, action, grounding or data
  • Ranked backlog with expected impact
  • A fixed test set for everything that follows

Topic & Instruction Refinement

Where classification is unstable or behaviour is off-policy. Boundaries redrawn against real phrasing, instructions rewritten as testable rules, and each change scored on its own.

  • Overlap analysis across live topics
  • Classification rewritten for contrast
  • Scope tightened or widened by evidence
  • Instructions made specific and testable
  • Customization & configuration →

Grounding & Action Quality

When the agent chose right and still answered badly. We test what retrieval returns per question and check whether actions succeed, time out or return something the agent cannot use.

  • Retrieved-passage review on failing questions
  • Content remediation where retrieval is faithful but wrong
  • Action success, failure and timeout rates
  • Input mapping and output shaping fixes
  • Data & knowledge integration →

Escalation & Handoff Redesign

Getting the moment right in both directions. Too eager destroys containment; too stubborn destroys satisfaction. Both are tuned against outcomes, not against a target number.

  • Escalation trigger review by topic
  • Failed-attempt and sentiment thresholds
  • Context handed to the human on transfer
  • Post-escalation outcome analysis
  • Queue routing and after-hours behaviour

Latency & Consumption Control

Cost and speed are mostly design outcomes. Fewer turns, fewer actions per turn, tighter responses and faster integrations all move both at once — without reducing what the agent can do.

  • Turn-count and action-count analysis per resolution
  • Slow action identification and timeout policy
  • Response length and verbosity tuning
  • Cost per resolved conversation, tracked
  • Test-run budget planning for future cycles

Observability & Review Cadence

The dashboards that make the next cycle possible, plus the meeting that uses them. Optimisation without a cadence is a one-off project that decays the month after it ends.

  • Dashboards segmented the way causes are
  • Alerting on containment and error thresholds
  • Monthly or quarterly review with a decision log
  • Regression suite re-run on every change
  • Testing & deployment →
What We Read

The signals a tuning cycle runs on

Optimisation is only as good as its instrumentation. These are the sources we connect, and what each one is uniquely able to tell you.

Session transcripts

The only source that shows why something went wrong rather than that it did. Reading two hundred real sessions produces a better backlog than any dashboard, and it is where the diagnostic always starts.

Agent observability reporting

Topic distribution, escalation rates and latency at the aggregate level. It tells you where to look and how big a problem is; it cannot tell you what caused it. Both halves matter.

Testing Center evaluations

Batch runs of a fixed test set, scoring expected topic and action against what actually happened. This is what makes “the change helped” a measurement rather than an impression — and it consumes budget, so runs are planned.

Case & CRM outcomes

Whether the case actually closed, whether it reopened a week later, whether the lead converted. Conversation-level metrics flatter an agent; record-level outcomes are where the truth is.

Consumption reporting

Usage per conversation and per topic, so cost can be attributed to design decisions rather than absorbed as a fixed platform expense. Without it, nobody can justify making the agent less chatty.

Human agent feedback

The people receiving escalations know which handoffs were premature and which arrived too late. A lightweight feedback loop from them is the cheapest signal in the whole programme, and the most ignored.

CSAT & survey data

The counterweight to containment. Tracked per topic rather than blended, so a well-contained but frustrating area cannot hide behind a healthy average across everything else.

The change log

What changed, when, and what it was expected to do. Without it, a metric that moves three weeks after four edits is unattributable — and the team relearns the same lesson every quarter.

How an Optimization Cycle Runs

One hypothesis, one change, one score

A first diagnostic runs two to three weeks. After that, tuning cycles are short and repeating — typically two to four weeks each, with a review at the end of every one.

Step 01

Instrument & Baseline

Segment the metrics that are currently blended, build the fixed test set from real sessions, and record where the agent stands today across every measure we intend to move.

Step 02

Diagnose to a Layer

Read sessions and attribute each failure to topic, action, grounding or data. The output is a ranked backlog where every item names its cause, not just its symptom.

Step 03

Change One Variable

Take the highest-expected-gain item, state what should improve and by roughly how much, and change only that — in a sandbox, with the current behaviour preserved.

Step 04

Re-Score & Regression Check

Run the same test set. Confirm the target measure moved and that previously-passing cases still pass — then promote, or discard the change and record why.

Step 05

Review & Repeat

Report what moved, log what changed, and start the next cycle from the updated backlog — with the test set growing as new failure patterns appear in production.

Client Outcomes

Measured improvement on systems already running

Diagnosing a live system, attributing causes and improving it against a baseline is the same discipline whether the system is an agent, a scoring model or a CRM.

★★★★★
The AI readiness engagement gave us a clear roadmap to operationalize AI across our processes. The team built intelligent ‘Next Best Action’ capabilities using scoring, engagement, and fit models, which significantly improved how we prioritize and interact with prospects. Their understanding of both CRM and AI-driven decisioning made a real difference in aligning our systems with business outcomes.
Rubesh J. Consultant · AI decisioning & prioritisation models Decisioning
Case Study

Professional Services Firm — AI Document Automation

Automating email attachment processing and Salesforce data routing.

40% Reduction in manual document processing time
2× Admin throughput without added headcount
100% Automated classification & CRM routing
Read Full Case Study
★★★★★
We engaged Twopir Consulting to conduct a Salesforce audit, and their structured, insight-driven approach exceeded our expectations. Their team quickly understood our complex processes, identified critical gaps, and provided clear, actionable recommendations. The audit improved our data accuracy and streamlined workflows.
Kelly Hale Managing Director · Diagnostic audit & remediation Diagnostic
Case Study

Professional Events Organization — CRM Integration

Salesforce–MeetMax integration for a corporate networking organization.

100% Elimination of manual data sync between platforms
360° Unified client view across events & CRM
0 Manual reconciliation tasks post-integration
Read Integration Story
Why Twopir

We will not tune what we have not measured

Optimisation is the easiest service to deliver badly, because a confident-sounding change and a real improvement look identical without a scoreboard.

A fixed test set, before anything changes

Built from real sessions and held stable across cycles. Without it, every claim about improvement is an opinion, and the agent's history becomes unknowable.

We diagnose to a layer before touching anything

Topic, action, grounding or data. Naming the layer prevents the most common waste in this work: rewriting instructions for a problem that lives in the knowledge base.

One variable per cycle

Slower in the short run and far faster over a quarter, because you learn which changes actually work instead of accumulating edits nobody can attribute.

We fix the cause, wherever it lives

If the answer is a content rewrite, a Flow change or a data resolution rather than an instruction, we do that work — because an optimisation practice that only touches prompts hits a ceiling fast.

Cost is a design metric, not a bill

We track consumption per resolved conversation and treat it as something the architecture controls, so the economics improve alongside the quality rather than against it.

Common Questions

Performance questions, answered with method

There is no industry number worth quoting you, because containment depends almost entirely on what mix of requests reaches the agent. An agent scoped to documented, transactional questions will contain a high share of them; the same agent exposed to everything, including requests it was never designed for, will look far worse on an identical build. That is why we measure containment per topic against your own baseline rather than against a benchmark — the only honest comparison is your agent last month.

Usually not. In our experience the majority of wrong answers trace to grounding — the agent retrieved a real passage that is stale, contradicted by another article, or written so ambiguously that it supports the wrong reading. The diagnostic step is simple: look at what was retrieved for the failing question. If the passage is wrong, no instruction rewrite will save it. If the passage was right and the answer was still wrong, then you are in instruction or topic territory.

Design, mostly. Whatever commercial model you are on, consumption follows how many turns a conversation takes and how many actions each turn triggers — so an agent that asks three clarifying questions before answering costs several times one that reads context it already has. The levers are shorter instructions, better input mapping so nothing is re-asked, fewer redundant actions per turn, and resolving conversations in fewer exchanges. Also budget for testing: evaluation runs consume capacity, so a heavy tuning programme has a real cost of its own.

The diagnostic takes two to three weeks and usually surfaces one or two changes that can ship immediately — a topic boundary, an over-cautious instruction, a badly mapped action input. Those show up in the numbers within days of promotion. Changes that depend on content remediation or data resolution take longer, because the fix is not in the agent at all. What we will not do is ship a batch of changes in week one and claim credit for whatever moves; attribution requires patience.

Because something around it changed. Agents depend on knowledge articles, records, Flows, integrations and product facts that all move independently — an article edited by a content owner, a new SKU with no documentation, a Flow updated by an admin, a policy changed by legal. The agent was correct at launch and is now quoting a world that no longer exists. This is why we treat a re-test cadence as part of the service rather than a nice-to-have: drift is the default state, not an incident.

Yes, and most should eventually. The method is learnable: keep a fixed test set, diagnose to a layer before changing anything, change one variable at a time, re-score, and log what you did. What a first engagement usually supplies is the diagnostic instinct — recognising that inconsistent answers mean topic overlap, that confident wrongness means grounding, that a latency spike means one slow action. We build the test set and the dashboards to be handed over, and several clients run their own cycles after two rounds with us.

Next Step

Show us a month of sessions. We will show you where the ceiling is

A performance diagnostic works from evidence you already have. We segment your metrics, read real sessions, attribute each failure to a layer, and hand you a ranked backlog with the expected gain from each item.

Salesforce, CRM & AI delivery for growing and mid-market companies · Contact the team