Salesforce · AI Agents

An agent that acts on your data is only as good as your data.

Agentforce agents do not just answer — they take action inside Salesforce, within guardrails you define. We design the topics, actions and grounding before anything is switched on, because a confident agent working from fragmented data is worse than no agent.

  • 2–6wk First agent live
  • 3–6mo Multi-agent programme
  • 24/7 Autonomous coverage
Request to Action
REQUESTS & GROUNDING Employees In the Salesforce UI · Slack Customers Portal · Messaging · Voice Knowledge articles CRM records Data 360 profiles AGENTFORCE LAYER · TWOPIR-ARCHITECTED Topics & Actions Scope · Guardrails Flow · Apex actions Prompt Templates Versioned · Grounded Structured output Trust Layer Sharing respected Masking · Audit trail AN AGENT ONLY SEES WHAT THE USER MAY SEE 2πr OUTCOMES Work Completed Records updated, not just described Deflected Volume Routine handled, complex escalated Governed AI Every interaction logged and auditable
12+
Years of Salesforce delivery
500+
Clients served worldwide
40+
Certified delivery specialists
98%
Client retention

Trusted by 500+ organizations — including RevOps, service and marketing operations teams putting AI agents into production on Salesforce with Twopir Consulting.

Social Justice Collaborative
Bernstein Liebhard LLP
LegalZoom
Sterling Law Offices
Sterling Law Offices
Social Justice Collaborative
Bernstein Liebhard LLP
LegalZoom
Sterling Law Offices
Sterling Law Offices

Built for Governed AI

  • Salesforce Partner
  • AI Delivery
  • Agentforce
  • Prompt Builder
  • Einstein Trust Layer
  • Agent Script
  • Data 360 Grounding
  • Agentforce Observability
Where Teams Stall

The real cost is operational friction

These are the patterns worth pointing an agent at first — repetitive, high-volume work where the answer already exists in Salesforce. They are also the ones where an agent is easiest to scope and easiest to prove.

Manual lead qualification slows everything down

Reps work through queues by hand while high-intent opportunities sit waiting. Scoring and prioritisation is exactly the kind of repeatable judgement an agent can carry, with the rep still making the call.

People spend their day searching for context

What someone needs is spread across accounts, opportunities, cases, emails and activity logs. Summarisation surfaces that context in place, instead of asking a person to reconstruct it from five tabs.

Repetitive queries consume the support team

Agents answer the same questions repeatedly, building backlog and leaving no capacity for the complex cases only a person can handle. This is the highest-volume, lowest-risk place to start.

Knowledge sits in Salesforce and never gets read

Articles, case histories and CRM records hold the answers but are slow to reach. An agent grounded in Knowledge surfaces them while still respecting your sharing model — nothing leaves the permission boundary.

Follow-ups are inconsistent because they are hand-written

Every rep drafts their own emails and case summaries with no shared structure. Versioned prompt templates produce consistent, reviewable output at scale rather than whatever each person types that day.

A pilot that impressed everyone and then stopped

The demo agent worked, the rollout stalled, and nobody can say whether it is helping. Almost always this is prompt quality, data completeness or adoption — and none of the three is visible without observability.

What It Is

Agents that take action, inside the permission model

Agentforce is the agent layer of the Salesforce platform: autonomous AI agents that work alongside your team and can act inside Salesforce rather than only answer questions. An agent is defined by topics — the areas it is allowed to handle — and actions, which are the Flows, Apex and standard operations it may invoke. Prompt Builder supplies versioned, grounded templates, and Agent Script pairs deterministic steps that must always run with the reasoning that handles the nuance in between.

The reason this is acceptable in a CRM at all is the Einstein Trust Layer: data retrieval that respects your sharing rules and field-level security, sensitive data masked before it reaches the model, no retention by the model provider, toxicity checks on responses, and an audit trail for every interaction. An agent cannot show a user something that user could not already open. That constraint is a feature, and it is what makes the difference between a governed deployment and a liability.

What Salesforce ships is the capability to build agents. What decides whether they are useful is the scope you give them and the data they are grounded on — which is the work Twopir Consulting is engaged to do. Grounding quality is usually the binding constraint, which is why this page and our Data 360 practice are two halves of the same conversation. See Salesforce services for how the layers fit together.

Do You Actually Need An Agent

Agent, Flow, or prompt template — they solve different problems

Plenty of work labelled "an AI use case" is a Flow with a good trigger. Agents cost more to run and more to govern, so they should be reserved for the problems that genuinely need reasoning — and we will tell you when yours does not.

Choosing between Flow, a prompt template and an Agentforce agent
CriteriaFlowPrompt TemplateAgentforce Agent
Use it whenThe rules are known and fixed, and the same input should always give the same output.You need generated language — a summary, a draft email — from data you already have.The path is not knowable in advance and the system must choose actions to reach a goal.
Who triggers itA record change, a schedule or a button.A user, or another process, on a specific record.A person in conversation, or an event, with the agent deciding what to do next.
PredictabilityDeterministic. Testable, and it behaves identically every time.Structured but generative — output varies within the shape you defined.Reasoned. Guardrails constrain it, but you are governing behaviour rather than a script.
Running costIncluded in the platform.Consumption on each generation.Consumption per action or conversation, plus Data 360 credits where grounded.
Governance effortLow — review the logic once.Moderate — version templates, review outputs.Highest — scope, guardrails, escalation paths and ongoing observability.
Our usual adviceStart here. Most "AI use cases" are a well-triggered Flow.Add when the bottleneck is genuinely writing, not deciding.Reserve for open-ended, high-volume work where reasoning earns its cost.
Three Different Engagements

Prove one agent, fix a stalled one, or build the hard ones

Three genuinely different jobs. On this product the first one should be small on purpose — a single agent, in production, with observability on it, is worth more than a six-agent roadmap nobody has validated.

Engagement 01

Assess & Prove

For teams with licences and no agent in production yet

We audit the org for AI readiness — data quality, object model, existing Flows, Knowledge coverage, sharing architecture — then design and ship one agent against the use case with the clearest return and the least risk.

  • AI readiness audit and prioritised use-case roadmap
  • Topic, action library and guardrail design
  • Prompt templates built, versioned and tested
  • Trust Layer configuration and escalation paths
  • Observability wired in before launch, not after
Where it stops

Delivered with standard actions, Flow and Prompt Builder. If an agent needs an Apex action or an external API call, that is scoped separately as build-on work — and sometimes the honest finding is that you do not need an agent for this at all.

Engagement 02

Rescue & Tune

For agents that are live but not trusted

The pilot impressed everyone and then plateaued. Answers are inconsistent, users have stopped asking, and nobody can say whether it is working. We instrument it, find where it is failing, and fix the cause rather than rewriting the prompt and hoping.

  • Observability build-out: sessions, deflection, latency, quality
  • Conversation log review against real failures
  • Prompt and instruction refinement, versioned
  • Topic scope and guardrail correction
  • Grounding gap analysis across Knowledge and CRM data
Where it stops

Tuning cannot fix a grounding problem. Where the answers are wrong because the underlying data is incomplete or fragmented, we say so and scope the data work — usually on Data 360 — rather than tightening prompts around a gap.

Engagement 03

Build On

For programmes past the first agent

Custom development for what configuration cannot reach: agents that call your own services, deterministic logic that must run in a fixed order, and the shared prompt and action libraries a multi-agent programme needs to stay coherent.

  • Apex actions and external API invocation
  • Agent Script for deterministic step sequencing
  • Shared action and prompt libraries across agents
  • Data 360 grounding and identity resolution design
  • Custom surfaces for agents outside standard channels
Where it stops

Every custom action widens what an agent can do, and therefore what it can get wrong. Each one is scoped with its own guardrails, failure behaviour and audit expectations before it is built — the boundary is argued in writing, not assumed.

What We Deliver

The agents we build, and the plumbing behind them

Each capability is scoped against a specific operational bottleneck, with a way to tell whether it worked — not switched on because the licence includes it.

Lead Qualification Agents

Scores and prioritises inbound leads against CRM data and engagement signals, then recommends the next action — so reps open the day with the right conversations rather than the most recent ones.

  • Scoring logic grounded in your own conversion history
  • Next-best-action recommendations on the record
  • Routing that respects existing assignment rules
  • Escalation to a human at defined thresholds
  • Measurement against pre-agent qualification rates

Grounded Summarisation & Drafting

Account, opportunity and case summaries, plus drafted follow-ups and case wrap-ups — generated from records the user is already permitted to see, in a structure you defined rather than whatever the model chooses.

  • Record summaries across standard and custom objects
  • Sales follow-up and outreach drafting
  • Case resolution summaries for clean reporting
  • Draft Knowledge articles from resolved cases
  • Output structure enforced by the template

Customer-Facing Service Agents

Agents on portals, messaging and voice that resolve routine volume around the clock and hand over to a person with the conversation already summarised — deflection that does not become a wall.

  • Conversation design and scope definition
  • Knowledge grounding with citation back to articles
  • Escalation with full context transferred
  • Deployment across portal, messaging and voice
  • Containment and satisfaction measured, not assumed

Actions, Agent Script & Orchestration

The part that turns an assistant into an agent: the actions it may invoke, the steps that must always run in order, and the guardrails around both.

  • Action library design across Flow and Apex
  • Agent Script for deterministic sequencing
  • Topic scoping so agents decline what is out of range
  • Multi-agent handoff where scope genuinely differs
  • Failure behaviour defined for every action

Trust Layer & Governance

Configured deliberately rather than accepted by default: what agents may retrieve, what gets masked, who can invoke which agent, and what the audit trail has to show your risk or compliance team.

  • Grounding scope aligned to the sharing model
  • Sensitive data masking configuration
  • Agent access by profile and permission set
  • Audit trail set up for the review you actually face
  • Documented policy for what agents may never do

Observability & Continuous Tuning

Wired in before launch, because without it you cannot tell a prompt problem from a data problem. Conversation logs, deflection, latency and quality scoring make refinement evidence-driven.

  • Session and conversation logging
  • Deflection and containment rate tracking
  • Latency monitoring on agent actions
  • Response quality scoring against real transcripts
  • Consumption tracking across both credit pools
Grounding & Integration

An agent is only as good as what it is allowed to read

Grounding is the difference between an agent that helps and one that invents. Each source below says what it contributes to an answer — and what it costs you to use it.

Salesforce Knowledge

The cheapest and most reliable grounding you have. Article quality becomes agent quality directly, which is why a Knowledge audit usually precedes any customer-facing agent.

Articles → agent · cited answers

CRM Records

Accounts, opportunities, cases and activity history, read strictly through the requesting user's permissions — the agent inherits their visibility, never more.

Records → agent · sharing enforced

Data 360

Unified profiles spanning systems the CRM alone cannot see. It raises answer quality materially — and meters on a separate credit pool, so it is a cost decision as well as a quality one.

Profiles → agent · separate credits

Flow & Apex Actions

What lets an agent do rather than describe. Existing Flows become actions without rebuilding them, and Apex covers what declarative logic cannot reach.

Agent → Salesforce · writes records

Experience Cloud Portals

Where a customer-facing agent usually lives. Portal content and permissions determine what it can answer, so the two are designed together rather than sequentially.

Portal ⇄ agent · authenticated context

Slack & Messaging Channels

Agents reachable where people already work, rather than behind another login. The same agent and the same guardrails, surfaced in a different place.

Chat ⇄ agent · same scope

External Systems via API

ERP, billing or product data an agent needs to answer properly, reached through governed actions with defined failure behaviour rather than an open call.

Agent ⇄ external · governed action

Observability & Digital Wallet

Conversation logs and quality scoring on one side, consumption across both credit pools on the other. Without both you are guessing at quality and at cost.

Agent → telemetry · logged per session
How We Deliver

Four phases, and an agent you can actually measure

A focused single-agent deployment — lead qualification, or an internal knowledge assistant — typically runs 2–6 weeks including discovery, build, testing and launch. A multi-agent programme involving Data 360, complex Flows and an org-wide prompt library runs 3–6 months.

Phase 01

Discovery & AI Readiness

We audit the org — data quality, object model, existing Flows, Knowledge coverage and sharing architecture — then map operational bottlenecks to specific capabilities and rank them by complexity against return. The output is a prioritised roadmap, including the cases where an agent is not the right answer.

Phase 02

Agent Design & Prompt Engineering

Topics, action libraries, Agent Script logic and prompt templates are designed before anything is built. Guardrails, escalation paths and Trust Layer configuration are settled first. Prompts get explicit instructions, grounding context and defined output structure, because that is what separates consistent behaviour from a good demo.

Phase 03

Build, Ground & Batch Test

We configure the agent, wire Flow and Apex actions, connect Knowledge and Data 360 sources, and test against real scenarios in volume rather than a handful of happy paths. Deployment is staged — a defined pilot group first, outputs validated, then wider rollout once behaviour holds.

Phase 04

Observability, Tuning & Adoption

After launch we use conversation logs, deflection, latency and quality scoring to find where instructions need sharpening and where scope needs tightening — and consumption tracking across both credit pools so cost never arrives as a surprise. Training and change management run alongside, because an unused agent is a failed one.

Common Scenarios

Where we see Agentforce deliver most

Four patterns that come up repeatedly. These are illustrative scenarios, not client case studies — we would rather show you the shape of the work than attach numbers we cannot attribute.

Lead Qualification at Inbound Scale

A company taking hundreds of inbound leads a week deploys a qualification agent. It evaluates each lead against CRM data and engagement signals, scores it, recommends the next action, and routes the high-priority ones to the right rep — before anyone reviews the queue by hand.

An Internal Knowledge Assistant

Reps and service agents ask for an account summary, a case history or a product detail and get an answer grounded in their own Salesforce records — respecting field-level security, so each person sees only what they are already permitted to open. No five tabs, no reconstructing context.

A Customer Agent on Self-Service

A public agent answers questions, explains services and guides prospects toward booking — around the clock, without a person involved in routine queries. When a question exceeds its defined scope it hands over to a human with the conversation already summarised.

AI-Assisted Case Handling

Support teams generate case summaries, surface relevant Knowledge articles and draft first responses inside the Service Cloud console. Cases close faster with more consistent quality, and the generated wrap-ups leave a clean trail for review — see our Service Cloud practice.

Why Twopir

Prompt quality, data completeness, and whether anyone uses it

Agentforce projects rarely stall at configuration. They stall where prompt quality, data completeness and user adoption meet — so we engineer all three before the first agent goes live, and measure all three after.

We will scope you out of an agent

If three Flows and a prompt template solve it, that is the recommendation — cheaper to run, easier to govern, and it leaves you able to add agents later on a foundation that supports them. We do not have a quota of agents to ship.

Prompt engineering treated as a discipline

Templates are versioned and tested the way code is, with explicit instructions, grounding context, defined output structure and edge-case handling. That is what produces predictable behaviour in production instead of demo-quality output that drifts.

Observability from day one, not after the complaint

Session tracking, deflection, latency and quality scoring are configured before launch. Without them you cannot separate a prompt problem from a data problem, and you will spend weeks rewriting instructions against a grounding gap.

We model both credit pools before you commit

Agentforce consumption and Data 360 consumption are metered separately, and the second is what makes a grounded agent cost more than the headline figure. We model both against your expected volume during design, not after the first invoice.

We know where the platform underneath will limit you

Twopir Consulting has delivered Salesforce for 12+ years across Sales, Service, Experience and data. That is how we can tell in an audit whether under-configured Flows, thin Knowledge or poor data quality will cap what an agent can do — before you build on it.

Common Questions

Answers before the first call

Agentforce is the agent layer of the Salesforce platform. It lets you build, deploy and govern autonomous AI agents that work alongside your team — answering questions, taking actions inside Salesforce, and serving customers across portals, messaging and voice. The distinction that matters is that an agent acts rather than only answers: it is defined by topics, which scope what it may handle, and actions, which are the Flows, Apex and standard operations it may invoke. Agents operate within guardrails you define and escalate to a person when a request falls outside them. Note that at Dreamforce 2025 Salesforce rebranded the Einstein 1 Platform as the Agentforce 360 Platform, so you will see both names in circulation.

Often a Flow would do the job, and it is worth checking before you spend anything. Use a Flow when the rules are known and fixed and the same input should always produce the same output — it is deterministic, testable, included in the platform and cheap to govern. Use a prompt template when you need generated language, such as a summary or a drafted email, from data you already hold. Reserve an agent for work where the path is not knowable in advance and the system has to choose actions to reach a goal. Agents cost more to run, because consumption is metered per action or conversation, and they cost more to govern, because you are managing behaviour rather than a script. A good assessment sometimes concludes that three Flows and one prompt template solve the problem, and that is a better outcome than an agent you have to supervise.

Through the Einstein Trust Layer, which is Salesforce's built-in governance framework for AI. It applies dynamic grounding, meaning data retrieval respects your existing sharing rules and field-level security; masking of sensitive data before it reaches the model; zero retention by the model provider, so your data is not used to train it; toxicity checks on responses; and an audit trail for every interaction. The practical consequence worth understanding is that an agent cannot surface anything the requesting user could not already open themselves. That constraint is the point — it means agent access is governed by the permission model you have already designed and tested, rather than by a separate one you would have to maintain alongside it.

Agentforce is consumption-priced rather than a flat per-seat licence, and there are two models: Flex Credits, drawn down per agent action, or per-conversation pricing. The important structural detail is that these are mutually exclusive within one org — you choose one consumption model, so it is an architecture decision rather than a procurement footnote. The second thing that catches teams out is that any agent grounded on Data 360 also consumes Data 360 credits from a separate pool, tracked alongside Agentforce usage. That means a Data 360-grounded agent costs more to run than the headline Agentforce figure implies, and both pools need modelling against expected volume before you commit to a design. Rates change, so confirm current pricing directly with Salesforce; what stays true is that consumption should be modelled during design, not discovered at renewal.

No. Agentforce works on standard Salesforce CRM data and Knowledge articles, and plenty of effective first agents are built without Data 360 at all. What Data 360 — formerly Data Cloud — adds is a unified customer profile spanning systems the CRM alone cannot see, which materially improves answer quality for anything that depends on a complete picture of the customer. The trade-off is cost, since Data 360 consumption is metered separately. A sensible sequence for most organisations is to prove one agent on CRM and Knowledge grounding first, measure whether the gaps in its answers are grounding gaps, and add Data 360 where the evidence says it will change the outcome rather than as a default starting assumption.

A focused single-agent deployment — a lead qualification agent, or an internal knowledge assistant — typically runs two to six weeks including discovery, build, testing and launch. A multi-agent enterprise programme involving Data 360, complex Flows and an org-wide prompt library runs three to six months. The variable that moves the timeline most is not the agent build itself but the state of what it grounds on: an org with well-maintained Knowledge and clean CRM data reaches production far faster than one where the first month is spent making the underlying data good enough to answer from.

Almost always one of three things, and they are hard to tell apart without instrumentation. Either the prompts are underspecified, so behaviour varies enough that users stop trusting it; or the grounding is incomplete, so the agent answers confidently from partial data and gets caught being wrong once, which is usually enough; or the agent was never scoped to work people actually do, so using it is more effort than not using it. The reason we configure observability before launch rather than after is that conversation logs, deflection rates and quality scoring separate these three cleanly. Without that evidence, teams tend to rewrite prompts repeatedly against what is actually a data problem — which is a slow and expensive way to discover the real cause.

Next Step

Start with readiness, not with a pilot

We assess the org an agent would have to run on — data quality, Knowledge coverage, sharing model, existing automation — and come back with the use cases worth building, the ones that are really a Flow, and what both credit pools would cost at your volume.

Response within 24 hours · We start with diagnosis, not a sales call · Contact the team