The action description was written for a developer
The reasoning engine reads that description to decide whether to call it. A terse technical label means the agent invokes the wrong tool, or never finds the right one.
Agentforce development is the engineering work behind an agent's capability: custom Apex and Flow actions, prompt templates, API-backed and MCP-exposed tools, and multi-agent designs. It starts exactly where declarative configuration stops — when the agent needs to reach a system, enforce a rule or complete a transaction that Agent Builder alone cannot express. Written to survive production traffic, not to pass a demo.
Trusted by 500+ organizations — including engineering teams who needed an agent to reach systems Salesforce does not own.
Development Scope
An agent action is called by a system that decides for itself when to use it, on inputs it assembles from a conversation. That is a harsher environment than a button on a page. Most of these failures are design failures, not coding mistakes.
The reasoning engine reads that description to decide whether to call it. A terse technical label means the agent invokes the wrong tool, or never finds the right one.
A three-second callout is fine in a batch job and painful mid-chat. Actions need a latency budget, a timeout, and a plan for what the agent says while waiting.
An action that swallows its exception returns nothing, and the agent improvises around the gap. Errors must come back as structured, describable outcomes.
Salesforce governor limits do not care that a human is waiting. Unbulkified Apex behind an agent action works in testing and fails on the day volume arrives.
Apex that runs without sharing gives an agent reach its user was never granted. That is a governance incident waiting for an auditor, not a shortcut.
Agent evaluation tests whether the agent chooses the action. It cannot tell you the action computes correctly. Both layers need their own tests.
Agentforce development is building the capabilities an agent can invoke. In Agentforce, an agent action is backed by metadata you write — an Apex class, a Flow or a prompt template — exposed to the reasoning engine through a name, a description and a typed set of inputs and outputs. Development means designing that interface so the agent selects it correctly, implementing the body so it behaves correctly under load and under least privilege, and defining what happens when it cannot do its job.
The interface is the part teams underestimate. Conventional Salesforce code is called by a button, a trigger or a schedule — something deterministic. An agent action is called by a reasoning engine that reads a natural-language description and decides. That makes the description part of the engineering, not documentation. An action named getAcctSts with no description will be chosen unpredictably no matter how good the Apex behind it is.
Keep the layers honest. Salesforce provides the action framework, Apex and Flow runtimes, prompt templates, the Trust Layer, and the governor limits everything runs inside. Twopir Consulting provides the design and the code: which capability should be an action at all, how its interface is described, bulk-safe implementation with sharing enforced, external calls with authentication, retries and timeouts, an error contract the agent can talk about, unit tests, and deployment as versioned metadata. You get capability that behaves the same on a bad Monday as it did in review.
We reach for code last, not first. If a Flow can express it, it should be a Flow — cheaper to change and legible to your admins. Our configuration practice exists partly to stop development engagements that should never have been sold. We work with growing, mid-market, and enterprise organizations that need help with complex CRM implementations, integrations, and business system challenges — and the cheapest code is still the code nobody had to write.
Picking the wrong backing technology is the most expensive early decision on an Agentforce build, because it is the hardest one to reverse later.
| Backed by | Choose it when | The trade-off |
|---|---|---|
| Flow | The work is record operations, branching or approvals your admins should be able to read and change. | Struggles with complex data shaping, loops at volume and elaborate error handling. |
| Apex | Logic is genuinely complex, needs bulk-safe processing, or must enforce rules a Flow cannot express. | Needs a developer to change, and needs real tests. Governor limits apply mid-conversation. |
| Prompt template | The output is generated text — a summary, a draft reply, a description — grounded in named fields. | Not deterministic. Never use one where an exact value or a transaction is required. |
| External API or MCP tool | The system of record sits outside Salesforce and the answer must be live rather than synced. | Latency, authentication and availability become yours to manage inside a conversation. |
If the agent behaves badly but can already reach what it needs, development is the wrong purchase. Declarative rework is faster and cheaper.
A capability gap, not a behaviour gap. The agent needs to reach a system, enforce a rule or complete a transaction it currently cannot.
Custom actions are rarely bought alone. If no agent exists yet, development runs inside a first implementation rather than beside it.
Six kinds of custom work. All of it is deployed as versioned metadata, reviewed like code, and handed over with tests rather than a walkthrough.
Invocable Apex exposed as an agent action, written to Salesforce engineering standards: bulk-safe, sharing-enforced, no hard-coded ids, and covered by tests that assert behaviour rather than lines.
with sharing and explicit CRUD and field checksWhere the logic is declarative, we build it as a Flow so your admins can maintain it. The engineering is in the input and output contract and in the failure paths, not in the clicks.
Templates that ground generation in named records and fields, fix the output structure, and fail predictably when the grounding is thin — instead of inventing something plausible.
Reaching systems Salesforce does not own. Named credentials, retry and timeout policy, response shaping, and — where MuleSoft is in play — exposing an existing API as a governed, agent-callable asset.
When one agent should not own everything, work is split across specialised agents with explicit handoffs — shared context, a clear owner per step, and an audit trail that survives the transfer.
Where your team writes the actions, we review them against the failure modes above and leave behind the patterns — so the second action does not repeat the first one's mistakes.
These are the integrations that come up repeatedly in development work, what moves in each direction, and the engineering constraint that decides the design.
The agent reads live order, stock and fulfilment state and writes back cancellations or address changes. The constraint is latency: an ERP that answers in four seconds needs an async pattern, not a longer wait.
Invoice state and payment history flow in as read-only context; anything that moves money goes out through an action with explicit confirmation and human approval. Idempotency is the engineering requirement here, not speed.
Rather than writing one Apex callout per system, an existing API is exposed as a governed asset the agent can call, carrying its own security policy, traffic control and tracing. Worth it once the second or third system appears.
The agent assembles the data, a generation service produces the document, and a signature request goes out with status flowing back to the record. The action returns a tracking reference, never a promise that it is done.
Where a model or ruleset outside Salesforce decides priority, eligibility or next best action, the agent calls it and uses the result as input to its own reasoning — with the score written to the record so the decision stays auditable.
Work too slow for a conversational turn is published as an event, processed asynchronously, and reported back later — so the agent can honestly say it has started something rather than blocking a customer for eight seconds.
Custom actions sometimes query the grounding layer directly for a structured lookup rather than a generated answer. Where the need is better retrieval rather than a new tool, that is data work, and we say so.
Agent metadata, Apex, Flows and prompt templates belong in the same repository and the same pipeline as everything else. Agents that are configured only by hand in production cannot be reviewed, diffed or rolled back.
Development runs in sprints against a defined action list. The first step is the one teams skip, and it is the one that decides whether the agent ever calls the thing correctly.
Name, description, inputs, outputs and the “when to use this” statement the reasoning engine reads — drafted and tested for selection accuracy before a line of logic is written.
Flow, Apex, prompt template or external call, decided against the table above rather than by habit — with the maintenance cost stated, not discovered.
Bulk-safe implementation with sharing enforced, a structured error contract, and unit tests that cover the failure paths as seriously as the success one.
Two separate questions: does the action compute correctly, and does the agent select and populate it correctly? Both are evaluated, because passing one proves nothing about the other.
Peer code review, deployment as versioned metadata through your pipeline, and handover with tests, documentation and the reasoning behind each design decision.
Automation, document intelligence and integration engineering are the same disciplines a custom agent action needs. These are Twopir builds where that engineering is the story.
Twopir helped us identify the right AI tools and select the best-fit platform for our needs. Their team provided end-to-end support — from consulting to implementation — delivering an AI-powered chatbot and automation system that improved our lead routing and customer engagement.
Automating email attachment processing and Salesforce data routing.
The AI readiness engagement gave us a clear roadmap to operationalize AI across our processes. The team built intelligent ‘Next Best Action’ capabilities using scoring, engagement, and fit models, which significantly improved how we prioritize and interact with prospects. Their understanding of both CRM and AI-driven decisioning made a real difference.
Salesforce–MeetMax integration for a corporate networking organization.
Custom agent work sits at an awkward intersection: it is Salesforce engineering, but the caller is non-deterministic. Teams that are strong at only one half produce actions that fail in predictable ways.
The action's name and description are what the reasoning engine uses to decide. We draft them first and test selection accuracy before implementing anything behind them.
Apex behind an agent action follows the same standards as any production Apex: bulkified, with sharing, explicit permission checks, no hard-coded ids, and tests that assert behaviour.
Every action has a defined contract for failure, so the agent can say what went wrong rather than apologising vaguely or claiming success it did not achieve.
If a Flow can do it, we build a Flow, because your admins can change it without us. A development engagement that shrinks during design is a good outcome, not a lost sale.
Agent configuration, Apex, Flows and templates live in source control and deploy like everything else — so changes are reviewable, diffable and reversible.
Use a Flow whenever a Flow can express it, because your admins can maintain it without a deployment. Reach for Apex when the logic involves genuinely complex data shaping, when it must process collections safely at volume, when it needs elaborate error handling or callout orchestration, or when it enforces a rule you do not want anyone changing by accident. The deciding question is not which is more powerful — it is who will need to change this in eighteen months, and what happens if they get it wrong.
Write the interface for the reader that actually uses it. The agent chooses an action by reading its name, description and input definitions, so those need to state plainly what the action does, when it should be used, and when it should not. Two actions with overlapping descriptions produce unstable selection in exactly the way two overlapping topics do. Attach actions only to the topics that need them, keep the input set small, and evaluate selection accuracy against real utterances before shipping.
Yes — through an action that makes the call, typically Apex or an external service using a named credential, or through MuleSoft where an existing API is exposed as a governed, agent-callable asset. The engineering considerations are latency inside a conversational turn, authentication and token handling, retry and timeout policy, and what the agent says when the external system is unavailable. Work that cannot complete inside a turn is better published as an asynchronous job the agent reports having started.
They inherit the permissions of the user the agent runs as, so object, field and record access all apply. The place that breaks is inside custom code: Apex declared without sharing runs in system context and will happily read records the agent user could never see. We write agent-facing Apex with sharing and add explicit object and field permission checks, then verify it as part of the release check rather than trusting the declaration. An action is the easiest place in an Agentforce build to accidentally grant more reach than anyone approved.
At two levels, because they answer different questions. Unit tests cover the action body — correct results, bulk behaviour, permission enforcement, and every failure path — and they are deterministic. Agent evaluation covers whether the agent selects the action for the right requests and populates its inputs correctly, which is not deterministic and needs a batch of real utterances scored against expected topic and action. An action can pass its unit tests perfectly and still never be called, and teams that test only one layer usually discover the other one in production.
Start with one, and split when you have a concrete reason. Good reasons to split: two domains with genuinely different risk profiles, different audiences such as customers versus employees, or a topic set that has grown large enough that classification is becoming unstable. Bad reasons: organisational boundaries, or an assumption that more agents means more capability. Every split adds a handoff, and handoffs are where context gets dropped — so the design work moves to what gets passed across and who owns the conversation afterwards.
Tell us what the agent needs to do and cannot. We will work through the interface, the backing technology, the failure modes and the security boundary — and tell you honestly if configuration would have solved it instead.
Salesforce, CRM & AI delivery for growing and mid-market companies · Contact the team