It answers a different question than the one asked
Classic topic overlap. Two classification descriptions both fit the request, so which one wins varies with phrasing — and looks random to everyone watching.
Agentforce customization and configuration is the declarative work inside Agent Builder — reshaping topics, rewriting instructions, rewiring actions, and changing how the agent behaves per channel — without writing code or re-running an implementation. Most agents that disappoint do not need rebuilding. They need their topic boundaries redrawn and their instructions written properly.
Trusted by 500+ organizations — including teams whose agents were technically live long before they were actually useful.
Configuration Scope
Every one of these is fixed declaratively, inside Agent Builder, without a developer or a deployment window. If your agent does any of them, the model is not the problem.
Classic topic overlap. Two classification descriptions both fit the request, so which one wins varies with phrasing — and looks random to everyone watching.
Three paragraphs where a sentence would do, or a chatty tone on a compliance topic. Length and register are instruction problems, and they are cheap to fix.
Scope was written as an aspiration rather than a boundary. An agent with no explicit refusals will try to be helpful about your refund policy, your legal position and your competitors.
Either it hands off the moment a question gets specific — so containment never moves — or it grinds through six failed turns before giving up, which is worse than no agent at all.
Action inputs are not mapped to available context, so the agent interrogates a logged-in customer for the order number sitting on the record in front of it.
The same agent serves an anonymous web visitor and a logged-in enterprise customer with the same assumptions. Channel-level configuration exists precisely to stop that.
Configuration is everything you can change about an agent's behaviour without writing code. In Agentforce that surface is larger than most teams realise. A topic carries a classification description that tells the reasoning engine when to select it, a scope statement that bounds what it may do, natural-language instructions that act as its operating rules, and example inputs that show it what real requests look like. Actions are attached to topics and wired to inputs. Prompt templates shape generated text. Channels decide where the agent appears and how it presents. All of that is declarative, all of it is reversible, and all of it changes behaviour immediately.
The distinction that matters commercially: configuration changes what the agent decides; development changes what the agent can reach. If your agent picks the wrong topic, rambles, refuses the wrong things, or escalates badly, that is configuration — days or weeks of work, no release pipeline. If it needs to call a system it currently cannot reach, or execute logic that no Flow can express, that is development. Partners who cannot draw that line tend to quote everything as a build.
Salesforce provides Agent Builder and the reasoning engine that acts on what you configure. Twopir Consulting provides the diagnosis and the craft: reading transcripts to find why a topic is being mis-selected, redrawing topic boundaries, and writing instructions that are specific enough to constrain behaviour without being so rigid that the agent stops reasoning. You get an agent that behaves the way your team would — measured against the same test set before and after, so the improvement is a number rather than an impression.
This is often the right first engagement for growing and mid-market companies that already bought Agentforce and are not getting value from it. It costs a fraction of a full implementation, and it tells you honestly whether the ceiling you have hit is a configuration ceiling or a data one.
This is the table we build during a configuration review, filled in from your own transcripts. Most rows resolve declaratively; the ones that do not are where a development conversation honestly begins.
| Symptom | The lever that fixes it | Code needed? |
|---|---|---|
| Picks the wrong topic | Rewrite classification descriptions so no two overlap; add real example utterances from transcripts. | No |
| Answers off-policy or too freely | Tighten topic scope and instructions; state refusals explicitly rather than implying them. | No |
| Re-asks for known information | Map action inputs to available conversation and record context instead of prompting for them. | No |
| Escalates at the wrong moment | Set explicit escalation conditions — failed attempts, sentiment, named intents — in instructions and channel setup. | No |
| Writes generated text you would not sign | Replace free generation with a prompt template that fixes structure, sources and tone. | No |
| Cites the wrong passage | A grounding problem, not a behaviour one: fix the data library, index and source content. | No — but it is data work |
| Cannot reach a system it needs | Requires a new action against an external API, with authentication and error handling. | Yes — development |
| Needs logic no Flow can express | Apex action, or an external service called through a governed API. | Yes — development |
Order matters. Rewriting instructions before fixing topic boundaries just makes a mis-classified request answer more fluently in the wrong place.
The first lever, because everything downstream depends on the right topic being selected. We redraw boundaries so each one owns a distinct job, then prove it against real phrasings.
Instructions are the agent's operating rules in plain language. We write them as specific, testable statements rather than the vague encouragement most agents ship with.
Connecting topics to the actions that do the work, and mapping inputs to context the agent already holds — so it stops interrogating people about things it can see.
Where the agent generates text — a summary, a reply, a description — a template fixes what it draws on and how the output is structured, which is the difference between useful and unpredictable.
The same agent should not behave identically to an anonymous visitor and a logged-in customer. We configure per-channel presentation, identity handling and handoff paths.
Configuration without measurement is redecorating. We build a test set from your transcripts, score the agent before we touch it, and score it again afterwards.
An agent's behaviour is shaped by settings that live in several places. A configuration engagement touches all of them, in both directions.
Deployment settings decide the greeting, the transfer target and what the human receives. Conversation and outcome flow back to the Case; the queue configuration flows the other way and decides who catches an escalation.
A Flow used as an agent action is configuration, not code. Its input and output variables define the contract with the agent; reshaping those is usually faster than changing how the agent asks.
Which library a topic can read is configuration; whether that library contains the right passage is content work. We separate the two on the first day so effort goes where the fault actually is.
An agent that says it cannot find a record is often permission-blocked rather than confused. Object and field access on the agent user constrain behaviour as firmly as any instruction does.
On a portal the site supplies identity and the sharing model supplies reach. Configuring the agent there means configuring both, because the same agent will behave differently for a guest and a member.
Masking and safety settings change what the model sees and what it is allowed to return. They are a governance decision with a direct behavioural effect, so we change them with the risk owner, not around them.
Topic distribution and escalation reporting are how you find the next thing to configure. They feed the backlog; the backlog feeds the next cycle. Without them configuration becomes guesswork.
Configuration is reversible, which tempts teams to edit production directly. We change in a sandbox, evaluate, then promote — because “reversible” and “harmless” are not the same word.
Two to six weeks, depending on how many topics are in play. The baseline in step two is what makes the last step meaningful — without it, nobody can say whether the work helped.
We go through real sessions, not a demo script, and classify what went wrong: mis-selected topic, thin grounding, bad instruction, missing action, or a genuine capability gap.
We build a test set from those transcripts and run it against the agent untouched, so every later claim of improvement has a number behind it.
Topic boundaries first, then instructions, then action wiring and templates — re-running the test set after each layer so we know which change produced which effect.
Per-channel behaviour, identity handling, escalation triggers and what the receiving human sees — the parts that decide whether a handoff feels competent.
Promote from sandbox with a regression run, then hand your admins a documented topic map, the test set, and the reasoning behind each instruction so they can keep it up.
A large share of our work is inherited systems that were technically live and operationally weak. These are engagements where the fix was configuration and structure, not a rebuild.
We engaged Twopir Consulting to conduct a Salesforce audit, and their structured, insight-driven approach exceeded our expectations. Their team quickly understood our complex processes, identified critical gaps, and provided clear, actionable recommendations. The audit improved our data accuracy, streamlined workflows, and aligned perfectly with our digital transformation goals.
Automating email attachment processing and Salesforce data routing.
The AI readiness engagement gave us a clear roadmap to operationalize AI across our processes. The team built intelligent ‘Next Best Action’ capabilities using scoring, engagement, and fit models, which significantly improved how we prioritize and interact with prospects. Their understanding of both CRM and AI-driven decisioning made a real difference.
Salesforce–MeetMax integration for a corporate networking organization.
A configuration engagement is smaller than an implementation and that is the point. If the fix is declarative, selling you a rebuild would be the easy thing and the wrong one.
The evidence for what is wrong is already in your session logs. Reading two hundred real conversations tells us more than any requirements meeting, and it costs you nothing to sit in.
Without a score from before the work, “it feels better” is the only available verdict. We measure first so the improvement is defensible to whoever approved the spend.
Topic boundaries before instructions, instructions before templates. Pulling levers out of order produces changes that cancel each other and a team that cannot tell what worked.
Sometimes the agent is fine and the knowledge behind it is wrong, or the capability genuinely needs building. We will say so and point at the right engagement rather than billing the wrong one.
We hand over a documented topic map, the test set and the reasoning behind each instruction. Configuration should be a skill your team owns, not a dependency you rent.
Look at what the agent gets wrong. If it reaches the right systems but chooses badly, says too much, refuses the wrong things or escalates at the wrong moment, that is configuration and it is measured in weeks. If it fundamentally cannot reach a system it needs, or the underlying process it is meant to support does not exist, no amount of configuration will help. A short transcript review settles it in days, and we would rather run that than quote a rebuild blind.
One job, described so distinctly that no other topic could plausibly claim the same request. A topic carries four things: a classification description that says when to select it, a scope statement that bounds what it may do, instructions that govern how it behaves, and example inputs drawn from real phrasing. The most common mistake is writing descriptions that are individually reasonable but collectively ambiguous — the test is not whether a human can tell them apart, it is whether they contrast sharply enough that classification is stable.
Specific, imperative and testable. “Be helpful and professional” changes nothing; “Answer in at most three sentences, quote the policy article you used, and if the customer asks about refunds outside the thirty-day window, transfer to the billing queue without offering an opinion” changes a great deal. Write refusals explicitly rather than hoping omission implies them, put verification steps before anything irreversible, and then test each instruction — if you cannot write a test case that would fail without it, it is decoration.
Yes, and they should own it long term — that is why we hand over the topic map, the test set and the reasoning behind each instruction. What a first engagement usually supplies is the diagnostic method rather than the clicks: knowing that inconsistent answers point at topic overlap rather than the model, that a chatty agent is an instruction problem, and that nothing should be judged without a before-and-after run. Most admin teams pick that up in one cycle and run the next one themselves.
It can, which is exactly why the test set covers paths that already work as well as the ones that fail. Changing a topic boundary to fix one misrouted intent can pull a neighbouring intent across with it, and you will not notice from spot checks. We work in a sandbox, run a regression after each layer of change, and promote only once the previously-passing cases still pass. Configuration being reversible is not the same as it being safe to do live.
Usually it is a grounding problem wearing a configuration costume. If the agent is retrieving a real passage from your content and that passage is out of date or contradicts another article, the agent is behaving correctly and the content is wrong. Configuration can reduce the blast radius — narrowing which library a topic reads, requiring a citation, instructing it to decline rather than infer — but the durable fix is in the knowledge layer. We check retrieval first before touching instructions, so effort lands where the fault is.
A configuration review starts with evidence you already have. We read real sessions, classify the failures, and come back with a list of levers in the order they should be pulled — and an honest note about anything that is not configuration at all.
Salesforce, CRM & AI delivery for growing and mid-market companies · Contact the team