Agentforce · AI Agent Development

When clicks run out, someone has to write the code.

Agentforce development is the engineering work behind an agent's capability: custom Apex and Flow actions, prompt templates, API-backed and MCP-exposed tools, and multi-agent designs. It starts exactly where declarative configuration stops — when the agent needs to reach a system, enforce a rule or complete a transaction that Agent Builder alone cannot express. Written to survive production traffic, not to pass a demo.

Anatomy of a Custom Action
WHAT THE AGENT CALLS Action Interface Label · Description · When to use Typed Contract Inputs · Outputs · Required flags Invoked Mid-Turn Under Agent User Latency Budget TWOPIR ENGINEERING LAYER Apex & Flow Bulk-safe · Governed Unit tested External Calls Auth · Retries · Timeouts Named credentials Error Contract Fail loud · Fail safe No silent success DEPLOYED AS METADATA · REVIEWED LIKE CODE 2πr WHAT SHIPS Reachable Systems the agent could not touch before Predictable Same inputs, same result, every time Maintainable Tested, documented and handed over DESIGN · BUILD · TEST · REVIEW · DEPLOY
12+
Years CRM & AI delivery
40+
Consultants & engineers
250+
Deployments delivered
15+
Technology partnerships

Trusted by 500+ organizations — including engineering teams who needed an agent to reach systems Salesforce does not own.

Amberscript
Kacific
Spinify
Zarraffa’s Coffee
Ultra Consultants
RChilli

Development Scope

  • Salesforce Partner
  • Apex Actions
  • Invocable Flows
  • Prompt Templates
  • REST & Named Credentials
  • MuleSoft & MCP
  • Platform Events
  • Salesforce DX & CI
Where Custom Work Goes Wrong

Custom actions that work in a demo and fail in the queue

An agent action is called by a system that decides for itself when to use it, on inputs it assembles from a conversation. That is a harsher environment than a button on a page. Most of these failures are design failures, not coding mistakes.

The action description was written for a developer

The reasoning engine reads that description to decide whether to call it. A terse technical label means the agent invokes the wrong tool, or never finds the right one.

It is too slow to sit inside a conversation

A three-second callout is fine in a batch job and painful mid-chat. Actions need a latency budget, a timeout, and a plan for what the agent says while waiting.

It fails silently and the agent apologises confidently

An action that swallows its exception returns nothing, and the agent improvises around the gap. Errors must come back as structured, describable outcomes.

It was written assuming one record at a time

Salesforce governor limits do not care that a human is waiting. Unbulkified Apex behind an agent action works in testing and fails on the day volume arrives.

It runs with more access than the conversation deserves

Apex that runs without sharing gives an agent reach its user was never granted. That is a governance incident waiting for an auditor, not a shortcut.

It has no tests, because “the agent tests it”

Agent evaluation tests whether the agent chooses the action. It cannot tell you the action computes correctly. Both layers need their own tests.

What It Means

What “custom Agentforce development” actually covers

Agentforce development is building the capabilities an agent can invoke. In Agentforce, an agent action is backed by metadata you write — an Apex class, a Flow or a prompt template — exposed to the reasoning engine through a name, a description and a typed set of inputs and outputs. Development means designing that interface so the agent selects it correctly, implementing the body so it behaves correctly under load and under least privilege, and defining what happens when it cannot do its job.

The interface is the part teams underestimate. Conventional Salesforce code is called by a button, a trigger or a schedule — something deterministic. An agent action is called by a reasoning engine that reads a natural-language description and decides. That makes the description part of the engineering, not documentation. An action named getAcctSts with no description will be chosen unpredictably no matter how good the Apex behind it is.

Keep the layers honest. Salesforce provides the action framework, Apex and Flow runtimes, prompt templates, the Trust Layer, and the governor limits everything runs inside. Twopir Consulting provides the design and the code: which capability should be an action at all, how its interface is described, bulk-safe implementation with sharing enforced, external calls with authentication, retries and timeouts, an error contract the agent can talk about, unit tests, and deployment as versioned metadata. You get capability that behaves the same on a bad Monday as it did in review.

We reach for code last, not first. If a Flow can express it, it should be a Flow — cheaper to change and legible to your admins. Our configuration practice exists partly to stop development engagements that should never have been sold. We work with growing, mid-market, and enterprise organizations that need help with complex CRM implementations, integrations, and business system challenges — and the cheapest code is still the code nobody had to write.

Choosing How to Build It

Four ways to back an action — and when each is right

Picking the wrong backing technology is the most expensive early decision on an Agentforce build, because it is the hardest one to reverse later.

Agentforce action types, when to choose each, and the trade-off it carries
Backed byChoose it whenThe trade-off
FlowThe work is record operations, branching or approvals your admins should be able to read and change.Struggles with complex data shaping, loops at volume and elaborate error handling.
ApexLogic is genuinely complex, needs bulk-safe processing, or must enforce rules a Flow cannot express.Needs a developer to change, and needs real tests. Governor limits apply mid-conversation.
Prompt templateThe output is generated text — a summary, a draft reply, a description — grounded in named fields.Not deterministic. Never use one where an exact value or a transaction is required.
External API or MCP toolThe system of record sits outside Salesforce and the answer must be live rather than synced.Latency, authentication and availability become yours to manage inside a conversation.
Default to the leftmost option that works. Every step right adds capability and adds maintenance.

Before this: Configure

If the agent behaves badly but can already reach what it needs, development is the wrong purchase. Declarative rework is faster and cheaper.

This page: Build on

A capability gap, not a behaviour gap. The agent needs to reach a system, enforce a rule or complete a transaction it currently cannot.

  • Custom Apex and Flow agent actions
  • External API and MCP-exposed tools
  • Prompt template engineering
  • Multi-agent and agent-to-agent patterns

Around this: Implement

Custom actions are rarely bought alone. If no agent exists yet, development runs inside a first implementation rather than beside it.

What We Build

Engineering work that an agent can be trusted with

Six kinds of custom work. All of it is deployed as versioned metadata, reviewed like code, and handed over with tests rather than a walkthrough.

Custom Apex Agent Actions

Invocable Apex exposed as an agent action, written to Salesforce engineering standards: bulk-safe, sharing-enforced, no hard-coded ids, and covered by tests that assert behaviour rather than lines.

  • Invocable methods with typed request and response
  • Bulkified processing and governor-limit safety
  • with sharing and explicit CRUD and field checks
  • Structured error returns the agent can describe
  • Unit tests covering success, failure and edge input

Flow-Based Agent Actions

Where the logic is declarative, we build it as a Flow so your admins can maintain it. The engineering is in the input and output contract and in the failure paths, not in the clicks.

  • Input and output variable design for agent use
  • Branching, approvals and record orchestration
  • Fault paths that return a describable outcome
  • Subflow structure for reuse across topics
  • Documentation your team can actually follow

Prompt Template Engineering

Templates that ground generation in named records and fields, fix the output structure, and fail predictably when the grounding is thin — instead of inventing something plausible.

  • Grounding field and related-record selection
  • Output structure, length and format constraints
  • Tone and register matched to channel
  • Explicit behaviour when data is missing
  • Side-by-side evaluation against human-written output

External API & MCP Tools

Reaching systems Salesforce does not own. Named credentials, retry and timeout policy, response shaping, and — where MuleSoft is in play — exposing an existing API as a governed, agent-callable asset.

  • Named credentials and external services
  • Timeout, retry and circuit-breaking policy
  • Response mapping into agent-usable shapes
  • MuleSoft MCP exposure of existing APIs
  • Latency measurement inside a conversation budget

Multi-Agent & Handoff Design

When one agent should not own everything, work is split across specialised agents with explicit handoffs — shared context, a clear owner per step, and an audit trail that survives the transfer.

  • Agent decomposition by domain and risk level
  • Context passed at handoff, not re-gathered
  • Ownership and escalation between agents
  • Loop and deadlock prevention
  • End-to-end tracing across the handoff

Developer Enablement & Review

Where your team writes the actions, we review them against the failure modes above and leave behind the patterns — so the second action does not repeat the first one's mistakes.

  • Code review against agent-specific failure modes
  • Reference implementations and action templates
  • Salesforce DX project and CI pipeline setup
  • Test strategy across Apex and agent evaluation
  • Testing & deployment →
What Custom Actions Reach

The systems that usually need code to get to

These are the integrations that come up repeatedly in development work, what moves in each direction, and the engineering constraint that decides the design.

ERP & order management

The agent reads live order, stock and fulfilment state and writes back cancellations or address changes. The constraint is latency: an ERP that answers in four seconds needs an async pattern, not a longer wait.

Billing & payments

Invoice state and payment history flow in as read-only context; anything that moves money goes out through an action with explicit confirmation and human approval. Idempotency is the engineering requirement here, not speed.

MuleSoft & the API estate

Rather than writing one Apex callout per system, an existing API is exposed as a governed asset the agent can call, carrying its own security policy, traffic control and tracing. Worth it once the second or third system appears.

Document generation & e-signature

The agent assembles the data, a generation service produces the document, and a signature request goes out with status flowing back to the record. The action returns a tracking reference, never a promise that it is done.

Scoring & decisioning services

Where a model or ruleset outside Salesforce decides priority, eligibility or next best action, the agent calls it and uses the result as input to its own reasoning — with the score written to the record so the decision stays auditable.

Platform events & async work

Work too slow for a conversational turn is published as an event, processed asynchronously, and reported back later — so the agent can honestly say it has started something rather than blocking a customer for eight seconds.

Data Cloud & retrieval

Custom actions sometimes query the grounding layer directly for a structured lookup rather than a generated answer. Where the need is better retrieval rather than a new tool, that is data work, and we say so.

Source control & CI

Agent metadata, Apex, Flows and prompt templates belong in the same repository and the same pipeline as everything else. Agents that are configured only by hand in production cannot be reviewed, diffed or rolled back.

How We Build

Interface first, implementation second

Development runs in sprints against a defined action list. The first step is the one teams skip, and it is the one that decides whether the agent ever calls the thing correctly.

Step 01

Design the Interface

Name, description, inputs, outputs and the “when to use this” statement the reasoning engine reads — drafted and tested for selection accuracy before a line of logic is written.

Step 02

Choose the Backing

Flow, Apex, prompt template or external call, decided against the table above rather than by habit — with the maintenance cost stated, not discovered.

Step 03

Build & Test the Body

Bulk-safe implementation with sharing enforced, a structured error contract, and unit tests that cover the failure paths as seriously as the success one.

Step 04

Test Through the Agent

Two separate questions: does the action compute correctly, and does the agent select and populate it correctly? Both are evaluated, because passing one proves nothing about the other.

Step 05

Review, Deploy, Hand Over

Peer code review, deployment as versioned metadata through your pipeline, and handover with tests, documentation and the reasoning behind each design decision.

Client Outcomes

Custom work that stayed built

Automation, document intelligence and integration engineering are the same disciplines a custom agent action needs. These are Twopir builds where that engineering is the story.

★★★★★
Twopir helped us identify the right AI tools and select the best-fit platform for our needs. Their team provided end-to-end support — from consulting to implementation — delivering an AI-powered chatbot and automation system that improved our lead routing and customer engagement.
Mathias Bensimon Director · AI platform selection & chatbot implementation AI Implementation
Case Study

Professional Services Firm — AI Document Automation

Automating email attachment processing and Salesforce data routing.

40% Reduction in manual document processing time
2× Admin throughput without added headcount
100% Automated classification & CRM routing
Read Full Case Study
★★★★★
The AI readiness engagement gave us a clear roadmap to operationalize AI across our processes. The team built intelligent ‘Next Best Action’ capabilities using scoring, engagement, and fit models, which significantly improved how we prioritize and interact with prospects. Their understanding of both CRM and AI-driven decisioning made a real difference.
Rubesh J. Consultant · AI readiness & Next Best Action engagement AI Ready
Case Study

Professional Events Organization — CRM Integration

Salesforce–MeetMax integration for a corporate networking organization.

100% Elimination of manual data sync between platforms
360° Unified client view across events & CRM
0 Manual reconciliation tasks post-integration
Read Integration Story
Why Twopir

Salesforce engineers who understand agents

Custom agent work sits at an awkward intersection: it is Salesforce engineering, but the caller is non-deterministic. Teams that are strong at only one half produce actions that fail in predictable ways.

We write the description as carefully as the code

The action's name and description are what the reasoning engine uses to decide. We draft them first and test selection accuracy before implementing anything behind them.

Bulk-safe and sharing-enforced by default

Apex behind an agent action follows the same standards as any production Apex: bulkified, with sharing, explicit permission checks, no hard-coded ids, and tests that assert behaviour.

Errors are designed, not discovered

Every action has a defined contract for failure, so the agent can say what went wrong rather than apologising vaguely or claiming success it did not achieve.

We argue for less code, not more

If a Flow can do it, we build a Flow, because your admins can change it without us. A development engagement that shrinks during design is a good outcome, not a lost sale.

It ships as metadata, through your pipeline

Agent configuration, Apex, Flows and templates live in source control and deploy like everything else — so changes are reviewable, diffable and reversible.

Common Questions

Development questions from technical teams

Use a Flow whenever a Flow can express it, because your admins can maintain it without a deployment. Reach for Apex when the logic involves genuinely complex data shaping, when it must process collections safely at volume, when it needs elaborate error handling or callout orchestration, or when it enforces a rule you do not want anyone changing by accident. The deciding question is not which is more powerful — it is who will need to change this in eighteen months, and what happens if they get it wrong.

Write the interface for the reader that actually uses it. The agent chooses an action by reading its name, description and input definitions, so those need to state plainly what the action does, when it should be used, and when it should not. Two actions with overlapping descriptions produce unstable selection in exactly the way two overlapping topics do. Attach actions only to the topics that need them, keep the input set small, and evaluate selection accuracy against real utterances before shipping.

Yes — through an action that makes the call, typically Apex or an external service using a named credential, or through MuleSoft where an existing API is exposed as a governed, agent-callable asset. The engineering considerations are latency inside a conversational turn, authentication and token handling, retry and timeout policy, and what the agent says when the external system is unavailable. Work that cannot complete inside a turn is better published as an asynchronous job the agent reports having started.

They inherit the permissions of the user the agent runs as, so object, field and record access all apply. The place that breaks is inside custom code: Apex declared without sharing runs in system context and will happily read records the agent user could never see. We write agent-facing Apex with sharing and add explicit object and field permission checks, then verify it as part of the release check rather than trusting the declaration. An action is the easiest place in an Agentforce build to accidentally grant more reach than anyone approved.

At two levels, because they answer different questions. Unit tests cover the action body — correct results, bulk behaviour, permission enforcement, and every failure path — and they are deterministic. Agent evaluation covers whether the agent selects the action for the right requests and populates its inputs correctly, which is not deterministic and needs a batch of real utterances scored against expected topic and action. An action can pass its unit tests perfectly and still never be called, and teams that test only one layer usually discover the other one in production.

Start with one, and split when you have a concrete reason. Good reasons to split: two domains with genuinely different risk profiles, different audiences such as customers versus employees, or a topic set that has grown large enough that classification is becoming unstable. Bad reasons: organisational boundaries, or an assumption that more agents means more capability. Every split adds a handoff, and handoffs are where context gets dropped — so the design work moves to what gets passed across and who owns the conversation afterwards.

Next Step

Bring us the capability your agent is missing, and we will design the action

Tell us what the agent needs to do and cannot. We will work through the interface, the backing technology, the failure modes and the security boundary — and tell you honestly if configuration would have solved it instead.

Salesforce, CRM & AI delivery for growing and mid-market companies · Contact the team