Salesforce · Litify AI & Agentforce

An agent reads the records your team creates. If those are thin, it will be confidently wrong.

Litify ships named agents for intake qualification, conflict checking, matter summaries, damages work and invoice review, and Salesforce Agentforce can act over the same data. All of them are only as good as what your team has been entering. We assess that honestly first, then configure agents with the guardrails and review steps the work actually requires. Readiness first. Agents second. We will tell you if you are not ready.

Agent Architecture
WHAT AGENTS READ Matter & Intake Records Case Type · Stage · Roles · Fields Documents & History Files · Chronology · Notes Time Entries Invoices Activity Timeline THE CONTROL LAYER Grounding What data an agent may actually see Guardrails What it may do, and what it may not Human Review Who checks it, and before what NO AGENT ACTS WITHOUT A REVIEW STEP 2πr WHAT MAKES IT SAFE A Scoped Remit Narrow enough to be checkable by a person An Audit Trail What it saw, did and who approved A Named Owner Someone accountable for the output
1
Readiness assessment before any agent goes live
14
Named LitifyAI capabilities to assess against
12+
Years of Salesforce delivery
AI Delivery
A dedicated delivery capability

Trusted by 500+ organizations — including law firms and legal technology companies building their case, billing and reporting operations on Salesforce with Twopir Consulting.

Social Justice Collaborative
Bernstein Liebhard LLP
LegalZoom
Sterling Law Offices, S.C.

AI Capability

  • Salesforce Partner
  • AI Delivery
  • Agentforce
  • LitifyAI Agents
  • Readiness Assessment
  • Guardrail Design
  • Human Review Workflows
  • Audit & Monitoring
Why AI Pilots Fail

Six reasons agents get switched off again

Almost none of these are failures of the model. They are failures of data, scope or accountability — which is why the fix is rarely a better prompt.

The records are too thin to reason over

An agent summarising a matter can only use what is on it. If the case type is generic, the stage was never updated and the notes live in someone's email, the summary will be fluent, plausible and wrong.

Inconsistent data makes the output unstable

Fourteen spellings of the same referral source, free text where a picklist belonged, stage values used differently by two teams. The agent reflects the inconsistency back, and nobody can tell whether it is the data or the model.

The remit is too broad to check

An agent asked to "help with intake" produces output nobody can verify. An agent asked to score an intake against six defined criteria produces output a person can check in thirty seconds.

No human review step anybody owns

Output that flows straight into a matter with no checkpoint. The first wrong answer that reaches a client is the last day that agent is enabled, and rightly so.

Nobody asked what the agent can see

Agents inherit the sharing model they are configured with. An agent that can read every matter in the firm is a confidentiality problem before it is an AI problem, and ethical walls do not enforce themselves.

No audit trail when it matters

A client or a regulator asks how a conclusion was reached. Without a record of what the agent saw, what it produced and who approved it, that question has no good answer.

What Is Actually On Offer

Two AI layers, and they do different jobs

A Litify org has two distinct AI surfaces, and confusing them is the most common reason firms scope this work badly. LitifyAI is the vendor's own set of named, purpose-built capabilities for legal work — agents for intake qualification, conflict checking, matter summaries, total policy, case value and strengths and weaknesses, plus a Damages Assistant, Instant Demands, a Document Assistant, AI time capture and AI invoice review. Litify ACE, its agentic case expert, extends this toward acting on a matter rather than only reporting on it.

Salesforce Agentforce is the platform's general agentic layer, available to any Salesforce org including one running Litify. It is configured rather than pre-built: you define topics, actions and the data it is grounded in. That makes it the right tool when the job is specific to your firm and no packaged agent fits. Litify positions the two as complementary, and in practice most firms end up using both.

What neither can do is compensate for weak data. Every one of these reads records your team creates — the case type, the stage, the matter fields, the documents, the time entries. That is why our first engagement is almost always a readiness assessment rather than a configuration, and why we are willing to recommend waiting. Vendor capability details are on Litify's own AI documentation.

Intake Qualifier

Scores and screens an intake against defined criteria as it is captured.

Conflict Check

Runs conflict checks from the matter, using fuzzy matching across case history.

Matter Summary

Produces a summary of a matter from the records and documents attached to it.

Damages & Demands

Reads records and bills into a source-linked chronology, then builds a demand packet.

AI Time Capture

Drafts time entries from activity, screened against client billing guidelines.

AI Invoice Review

Reviews inbound outside-counsel invoices line by line against billing rules and budgets.

Litify ACE

The agentic case expert — oriented toward driving work forward, not just reporting on it.

Agentforce

The platform layer: your own topics and actions, grounded in your own data.

Three Stages

Assess, pilot, then scale — in that order

We have not yet seen a firm benefit from skipping the first stage. The assessment is cheap and it frequently changes what the other two should contain.

Stage 01 · Assess

Can your data support an agent at all?

A readiness assessment against the specific capabilities you are considering. Different agents need different data to be good, so this is assessed per agent rather than as one score.

  • Data completeness on the fields each agent reads
  • Case type and stage consistency measured
  • Document coverage and classification quality
  • Sharing model implications for agent access
  • Which agents are viable now, and which are not
  • What would have to change to make the rest viable

Honest outcome Sometimes: not yet. We would rather say that than configure an agent that produces confident output nobody should act on.

Stage 02 · Pilot

One agent, one team, one measurable claim

A narrow deployment with a scoped remit, a human review step and a measurable before-and-after. Narrow enough that a person can check every output during the pilot.

  • One agent, one practice area, one team
  • A remit narrow enough to verify by hand
  • Human review on every output initially
  • Guardrails and escalation defined up front
  • A baseline measured before switch-on
  • An exit criterion agreed in advance

Honest outcome A pilot that does not meet its criterion gets switched off. That is a successful pilot — it cost you weeks rather than a firm-wide rollout.

Stage 03 · Scale

More agents, or wider deployment

Only once a pilot has earned it. Scaling changes the risk profile: more matters, more users, less individual scrutiny, and the review step has to be redesigned for that.

  • Review step redesigned for volume
  • Sampling rather than full manual review
  • Monitoring and drift detection in place
  • Audit trail retention agreed
  • Additional agents sequenced, not batched
  • A named owner for each agent in production

Honest outcome Scaling reduces per-output scrutiny by design. If you cannot describe how errors get caught at volume, you are not ready to scale.

What each AI capability reads, and what has to be true in your data before it is worth enabling
CapabilityWhat it reads, and what must be true firstReadiness
Intake qualificationThe intake record and questionnaire answers. Needs a consistent case type taxonomy and questionnaires that actually branch — a generic intake form gives it nothing to score.Data-light
Conflict checkingClient and party records across case history. Needs reasonable name and party data quality; fuzzy matching helps with typos but not with parties never recorded as Roles.Data-light
AI time captureActivity, tasks, emails and calendar events. Needs activity actually captured in the system rather than work happening outside it.Data-medium
Matter summarisationThe whole matter: fields, stage history, documents, notes and activity. Thin matters produce fluent, plausible and wrong summaries — this is the one most often enabled too early.Data-heavy
Damages and demand packetsMedical records and bills attached to the matter. Needs documents present, classified and legible; scanned files that were never OCR'd give it nothing.Data-heavy
Invoice reviewInbound invoices against billing rules and budgets. Needs the rules and budgets to exist as data, not as a PDF somebody emailed.Data-heavy
What We Deliver

Six workstreams, and readiness is always the first

The order is not negotiable. Everything after the assessment depends on what it finds, which is why we will not quote a configuration before doing one.

AI Readiness Assessment

Measured per capability, because different agents need different data to be good. The output names which agents are viable now, which are not, and what would change that.

  • Field completeness on what each agent reads
  • Case type and stage consistency measured
  • Document coverage and classification quality
  • Activity capture rates by role
  • Sharing model implications for agent access
  • A viable-now list and a not-yet list

LitifyAI Agent Enablement

Configuring and tuning the vendor's named agents for your practice — the fastest route to value where a packaged agent genuinely fits the job.

  • Capability selection against your actual work
  • Configuration and tuning per practice area
  • Scoped remits narrow enough to verify
  • Output review workflows
  • Measurement against a real baseline
  • Sequenced rollout, not a batch switch-on

Agentforce Topics & Actions

Where no packaged agent fits. Custom topics and actions grounded in your Litify data, built with the same guardrails as everything else.

  • Topic design against a defined job
  • Actions scoped to what it may actually do
  • Grounding on the right data, and only that
  • Prompt and instruction design
  • Testing against real matter scenarios
  • Fallback and escalation to a person

Guardrails & Access Design

What an agent may see and do. This is a confidentiality question before it is an AI question, and ethical walls do not enforce themselves.

  • Agent access against the org sharing model
  • Ethical wall and matter isolation respected
  • Actions the agent may not take, enumerated
  • Confidence thresholds and escalation
  • Client data handling and retention
  • What is logged, and for how long

Human Review Workflows

The checkpoint that makes output safe to act on — designed so it is genuinely used rather than clicked through.

  • Review step placed before consequence
  • Reviewer role and accountability named
  • What a reviewer is actually checking
  • Rejection and correction paths
  • Sampling design once volume grows
  • Feedback loop back into tuning

Monitoring, Audit & Drift

What you need when a client or a regulator asks how a conclusion was reached — and what tells you the agent has quietly got worse.

  • Audit trail of input, output and approver
  • Output quality sampled over time
  • Drift detection as data and practice change
  • Usage and override-rate reporting
  • A named owner per agent in production
  • A documented switch-off procedure
Our Position

Six things we will say that a vendor will not

We implement this technology and we think it is genuinely useful. That is exactly why we are direct about the conditions under which it is not.

Thin Data Beats No Agent — Rarely

An agent reading sparse matter records does not produce a cautious answer. It produces a fluent, plausible and wrong one, which is considerably more dangerous than no answer at all.

So we Assess data before configuring, per capability, and name the agents that are not viable yet.

Narrow Remits Beat Clever Ones

An agent scoring an intake against six defined criteria is checkable in thirty seconds. An agent "helping with intake" produces output nobody can verify, and unverifiable output does not get trusted.

So we Scope every agent to a job a person could check by hand, especially during a pilot.

A Review Step Is Not Optional

Output that flows straight into a matter with no checkpoint is a question of when, not whether. The first wrong answer that reaches a client is the last day that agent runs.

So we Place a review step before any consequence, with a named accountable reviewer.

Agent Access Is A Confidentiality Question

An agent that can read every matter in the firm is a problem regardless of how well it performs. Ethical walls and matter isolation have to be designed into what the agent can see.

So we Design agent access against the org sharing model before any agent is enabled.

Measure Against A Real Baseline

"It feels faster" is not a result. Without a measured before, you cannot tell whether an agent helped, and you certainly cannot defend the spend at a partners' meeting.

So we Measure a baseline before switch-on and agree the exit criterion in advance.

Sequence After Go-Live, Usually

During an implementation the data is at its thinnest and the team is at its busiest. Agents configured then read records that do not yet reflect how the firm will actually work.

So we Normally recommend agents as the phase after go-live, once the data they depend on is real.
How We Deliver

Five phases, and one of them can stop the project

The assessment has a genuine veto. An engagement that concludes "not yet, and here is why" is a successful one — it has saved you a rollout that would have been withdrawn.

Phase 01

Assess Readiness

Per capability, against the data each agent actually reads. Field completeness, case type and stage consistency, document coverage, activity capture. Output: a viable-now list and a not-yet list.

Phase 02

Design Guardrails

What the agent may see, what it may do, what it may never do, where the review step sits and who owns it. Designed against the org sharing model, not bolted on afterwards.

Phase 03

Configure & Test

LitifyAI capabilities tuned, or Agentforce topics and actions built. Tested against real matter scenarios from your own history, including the awkward ones.

Phase 04

Pilot & Measure

One agent, one team, a measured baseline and an exit criterion agreed in advance. Human review on every output while the pilot runs.

Phase 05

Scale Or Stop

If the criterion is met, scale with a review step redesigned for volume and monitoring in place. If it is not, we switch it off and tell you why.

Why the assessment can stop everything Agents read the records your team creates. If case types are inconsistent, stages were never reliably updated or documents were never classified, an agent will produce output that is fluent, confident and wrong — and the firm will conclude that AI does not work, when what did not work was the data underneath it. Fixing that first is usually cheaper than a withdrawn rollout, and it improves reporting and operations whether or not you ever enable an agent.

Where Agents Earn Their Place

Four jobs worth starting with

These have the best ratio of value to risk in most firms: narrow enough to verify, with data that usually already exists.

Intake Qualification

Scoring intakes against defined criteria as they arrive. Narrow, checkable, and the data is captured at the moment of entry rather than depending on later discipline.

Conflict Checking

Fuzzy matching across case history catches what an exact-match search misses. Low risk because the output is a list a person reviews, not an action taken.

Time Capture Assistance

Drafting entries from activity, screened against billing guidelines. High value on realization, and the timekeeper reviews every entry before it is submitted.

Invoice Review

For in-house teams reviewing outside counsel bills line by line against rules and budgets. Works well because the rules are explicit and the output is a flag, not a decision.

Client Outcomes

Legal platforms we have actually built

Two engagements from our legal practice. Both produced the clean, consistent operating data that any agent depends on — which is the work that has to come first.

★★★★★
Twopir provided Salesforce customisation and integration services to help us build a robust, compliant, and scalable legal operations platform — connecting case management, document processing, and financial systems into one unified workflow. The result was transformative for how we run case-to-cash operations.
Operations Lead Fast-growing personal injury law firm Personal Injury
Case Study

Personal Injury Firm — Multi-State

Streamlining case-to-cash operations with Salesforce, AWS and QuickBooks.

40%+ Faster case-to-settlement processing
45% Reduction in reconciliation effort
35% Improvement in data accuracy
Read Full Case Study
★★★★★
Twopir's specialized Salesforce customization enabled efficient integration of third-party systems and streamlined administration and billing, leading to seamless financial operations and enhanced productivity. Automated mass billing and matter management minimized errors across our entire legal workflow.
Practice Manager Mid-size US family law firm · 150 employees Family Law
Case Study

Family Law Firm — 150 Employees, US

A 50% efficiency gain from Accounting Seed and Salesforce integration.

50% Increase in operational efficiency
45% Productivity gains from automation
35% Faster lead qualification & conversion
Read Integration Story
Why Twopir

The partner willing to say "not yet"

Everyone in this market will configure agents for you. The useful question is whether they will tell you when your data cannot support them.

We assess readiness per agent, not as one score

Intake qualification needs very different data from matter summarisation. A single "AI readiness score" is marketing; a per-capability assessment tells you which agents to enable this quarter and which to wait on.

We will recommend waiting

And we have. A withdrawn rollout costs more than a delayed one, and it teaches a firm that AI does not work when what did not work was the data underneath it.

Guardrails before capability

What an agent may see and do is a confidentiality question before it is an AI one. Ethical walls and matter isolation get designed into agent access, not assumed.

Every pilot has an exit criterion agreed in advance

Measured baseline, defined success, and a switch-off decision that was made before anyone became attached to the project. "It feels faster" is not a result.

We work with growing and mid-market companies

We help growing and mid-market companies solve complex CRM, integration and business system challenges, and we serve enterprise organizations with the same architecture discipline. Firms at that stage need a system that survives the next three years of growth — not one built for the org chart they had last year.

Common Questions

Answers before the first call

LitifyAI is the vendor's own set of purpose-built legal capabilities — named agents for intake qualification, conflict checking, matter summaries, total policy, case value and strengths and weaknesses, plus a Damages Assistant, Instant Demands, a Document Assistant, AI time capture and AI invoice review. Litify ACE, its agentic case expert, extends this toward driving work forward rather than only reporting. Agentforce is Salesforce's general agentic layer, available to any org: you define topics, actions and grounding data yourself. Packaged agents are faster where one fits your job; Agentforce is the answer where nothing fits. Most firms end up using both.

Usually not in the same phase, and this is one of the few things we are firm about. During an implementation the data is at its thinnest and the team is at its busiest, so agents configured then read records that do not yet reflect how the firm will actually work. We assess readiness during discovery and normally sequence agents into the phase after go-live, once the case types are settled, stages are being updated reliably and the documents are where they should be. The data work that makes agents viable improves reporting and operations regardless.

That is what a readiness assessment measures, and it is measured per capability rather than as a single score — intake qualification needs very different data from matter summarisation. We look at field completeness on exactly what each agent reads, case type and stage consistency across the record set, document coverage and classification quality, and activity capture rates by role. The output is a viable-now list, a not-yet list, and a specific description of what would have to change to move an agent from the second list to the first.

The org's sharing model, if it is designed to. Agents operate within the access their configuration grants, so this is a confidentiality question before it is an AI question — an agent that can read every matter in the firm is a problem regardless of how well it performs. We design agent access explicitly against the sharing model, including ethical walls and matter isolation, and we enumerate the actions an agent may never take. This is part of the guardrail phase, before any agent is enabled, rather than something checked afterwards.

A human review step catches it, which is why we place one before any consequence and name an accountable reviewer rather than leaving it to whoever is nearest. During a pilot every output is reviewed; at scale that becomes sampling, with monitoring on override rates and output quality over time. Every agent action is logged with what it saw, what it produced and who approved it — which is what you need when a client or a regulator asks how a conclusion was reached. Every agent also has a documented switch-off procedure and a named owner.

Only if a baseline is measured before switch-on, which is why we insist on it. "It feels faster" is not a result and will not survive a partners' meeting. We agree a measurable claim and an exit criterion in advance — time to qualify an intake, time to draft a demand, override rate on suggested time entries — and measure against the baseline at the end of the pilot. If the criterion is not met we switch the agent off and tell you why. That is a successful pilot: it cost weeks rather than a firm-wide rollout.

Sometimes, and the honest answer depends more on your data than on your size. The capabilities with the best value-to-risk ratio for most firms are narrow and checkable: intake qualification, conflict checking, time capture assistance and invoice review. They work because the remit is tight, the output is reviewable in seconds, and the data they need is usually already being captured. Matter summarisation and damages work need much richer records and are where firms most often enable too early. We will tell you which group you are in.

Next Step

Before you turn agents on, find out what they will be reading

A readiness assessment measures whether your data can support the specific capabilities you are considering, and names the ones that are viable now. If the answer is not yet, you will get that answer with the reasoning.

Readiness per capability · guardrails · scoped pilot · measured against a baseline