The records are too thin to reason over
An agent summarising a matter can only use what is on it. If the case type is generic, the stage was never updated and the notes live in someone's email, the summary will be fluent, plausible and wrong.
Litify ships named agents for intake qualification, conflict checking, matter summaries, damages work and invoice review, and Salesforce Agentforce can act over the same data. All of them are only as good as what your team has been entering. We assess that honestly first, then configure agents with the guardrails and review steps the work actually requires. Readiness first. Agents second. We will tell you if you are not ready.
Trusted by 500+ organizations — including law firms and legal technology companies building their case, billing and reporting operations on Salesforce with Twopir Consulting.








AI Capability
Almost none of these are failures of the model. They are failures of data, scope or accountability — which is why the fix is rarely a better prompt.
An agent summarising a matter can only use what is on it. If the case type is generic, the stage was never updated and the notes live in someone's email, the summary will be fluent, plausible and wrong.
Fourteen spellings of the same referral source, free text where a picklist belonged, stage values used differently by two teams. The agent reflects the inconsistency back, and nobody can tell whether it is the data or the model.
An agent asked to "help with intake" produces output nobody can verify. An agent asked to score an intake against six defined criteria produces output a person can check in thirty seconds.
Output that flows straight into a matter with no checkpoint. The first wrong answer that reaches a client is the last day that agent is enabled, and rightly so.
Agents inherit the sharing model they are configured with. An agent that can read every matter in the firm is a confidentiality problem before it is an AI problem, and ethical walls do not enforce themselves.
A client or a regulator asks how a conclusion was reached. Without a record of what the agent saw, what it produced and who approved it, that question has no good answer.
A Litify org has two distinct AI surfaces, and confusing them is the most common reason firms scope this work badly. LitifyAI is the vendor's own set of named, purpose-built capabilities for legal work — agents for intake qualification, conflict checking, matter summaries, total policy, case value and strengths and weaknesses, plus a Damages Assistant, Instant Demands, a Document Assistant, AI time capture and AI invoice review. Litify ACE, its agentic case expert, extends this toward acting on a matter rather than only reporting on it.
Salesforce Agentforce is the platform's general agentic layer, available to any Salesforce org including one running Litify. It is configured rather than pre-built: you define topics, actions and the data it is grounded in. That makes it the right tool when the job is specific to your firm and no packaged agent fits. Litify positions the two as complementary, and in practice most firms end up using both.
What neither can do is compensate for weak data. Every one of these reads records your team creates — the case type, the stage, the matter fields, the documents, the time entries. That is why our first engagement is almost always a readiness assessment rather than a configuration, and why we are willing to recommend waiting. Vendor capability details are on Litify's own AI documentation.
Scores and screens an intake against defined criteria as it is captured.
Runs conflict checks from the matter, using fuzzy matching across case history.
Produces a summary of a matter from the records and documents attached to it.
Reads records and bills into a source-linked chronology, then builds a demand packet.
Drafts time entries from activity, screened against client billing guidelines.
Reviews inbound outside-counsel invoices line by line against billing rules and budgets.
The agentic case expert — oriented toward driving work forward, not just reporting on it.
The platform layer: your own topics and actions, grounded in your own data.
We have not yet seen a firm benefit from skipping the first stage. The assessment is cheap and it frequently changes what the other two should contain.
A readiness assessment against the specific capabilities you are considering. Different agents need different data to be good, so this is assessed per agent rather than as one score.
Honest outcome Sometimes: not yet. We would rather say that than configure an agent that produces confident output nobody should act on.
A narrow deployment with a scoped remit, a human review step and a measurable before-and-after. Narrow enough that a person can check every output during the pilot.
Honest outcome A pilot that does not meet its criterion gets switched off. That is a successful pilot — it cost you weeks rather than a firm-wide rollout.
Only once a pilot has earned it. Scaling changes the risk profile: more matters, more users, less individual scrutiny, and the review step has to be redesigned for that.
Honest outcome Scaling reduces per-output scrutiny by design. If you cannot describe how errors get caught at volume, you are not ready to scale.
Scroll the table sideways →
| Capability | What it reads, and what must be true first | Readiness |
|---|---|---|
| Intake qualification | The intake record and questionnaire answers. Needs a consistent case type taxonomy and questionnaires that actually branch — a generic intake form gives it nothing to score. | Data-light |
| Conflict checking | Client and party records across case history. Needs reasonable name and party data quality; fuzzy matching helps with typos but not with parties never recorded as Roles. | Data-light |
| AI time capture | Activity, tasks, emails and calendar events. Needs activity actually captured in the system rather than work happening outside it. | Data-medium |
| Matter summarisation | The whole matter: fields, stage history, documents, notes and activity. Thin matters produce fluent, plausible and wrong summaries — this is the one most often enabled too early. | Data-heavy |
| Damages and demand packets | Medical records and bills attached to the matter. Needs documents present, classified and legible; scanned files that were never OCR'd give it nothing. | Data-heavy |
| Invoice review | Inbound invoices against billing rules and budgets. Needs the rules and budgets to exist as data, not as a PDF somebody emailed. | Data-heavy |
The order is not negotiable. Everything after the assessment depends on what it finds, which is why we will not quote a configuration before doing one.
Measured per capability, because different agents need different data to be good. The output names which agents are viable now, which are not, and what would change that.
Configuring and tuning the vendor's named agents for your practice — the fastest route to value where a packaged agent genuinely fits the job.
Where no packaged agent fits. Custom topics and actions grounded in your Litify data, built with the same guardrails as everything else.
What an agent may see and do. This is a confidentiality question before it is an AI question, and ethical walls do not enforce themselves.
The checkpoint that makes output safe to act on — designed so it is genuinely used rather than clicked through.
What you need when a client or a regulator asks how a conclusion was reached — and what tells you the agent has quietly got worse.
We implement this technology and we think it is genuinely useful. That is exactly why we are direct about the conditions under which it is not.
An agent reading sparse matter records does not produce a cautious answer. It produces a fluent, plausible and wrong one, which is considerably more dangerous than no answer at all.
So we Assess data before configuring, per capability, and name the agents that are not viable yet.An agent scoring an intake against six defined criteria is checkable in thirty seconds. An agent "helping with intake" produces output nobody can verify, and unverifiable output does not get trusted.
So we Scope every agent to a job a person could check by hand, especially during a pilot.Output that flows straight into a matter with no checkpoint is a question of when, not whether. The first wrong answer that reaches a client is the last day that agent runs.
So we Place a review step before any consequence, with a named accountable reviewer.An agent that can read every matter in the firm is a problem regardless of how well it performs. Ethical walls and matter isolation have to be designed into what the agent can see.
So we Design agent access against the org sharing model before any agent is enabled."It feels faster" is not a result. Without a measured before, you cannot tell whether an agent helped, and you certainly cannot defend the spend at a partners' meeting.
So we Measure a baseline before switch-on and agree the exit criterion in advance.During an implementation the data is at its thinnest and the team is at its busiest. Agents configured then read records that do not yet reflect how the firm will actually work.
So we Normally recommend agents as the phase after go-live, once the data they depend on is real.The assessment has a genuine veto. An engagement that concludes "not yet, and here is why" is a successful one — it has saved you a rollout that would have been withdrawn.
Per capability, against the data each agent actually reads. Field completeness, case type and stage consistency, document coverage, activity capture. Output: a viable-now list and a not-yet list.
What the agent may see, what it may do, what it may never do, where the review step sits and who owns it. Designed against the org sharing model, not bolted on afterwards.
LitifyAI capabilities tuned, or Agentforce topics and actions built. Tested against real matter scenarios from your own history, including the awkward ones.
One agent, one team, a measured baseline and an exit criterion agreed in advance. Human review on every output while the pilot runs.
If the criterion is met, scale with a review step redesigned for volume and monitoring in place. If it is not, we switch it off and tell you why.
Why the assessment can stop everything Agents read the records your team creates. If case types are inconsistent, stages were never reliably updated or documents were never classified, an agent will produce output that is fluent, confident and wrong — and the firm will conclude that AI does not work, when what did not work was the data underneath it. Fixing that first is usually cheaper than a withdrawn rollout, and it improves reporting and operations whether or not you ever enable an agent.
These have the best ratio of value to risk in most firms: narrow enough to verify, with data that usually already exists.
Scoring intakes against defined criteria as they arrive. Narrow, checkable, and the data is captured at the moment of entry rather than depending on later discipline.
Fuzzy matching across case history catches what an exact-match search misses. Low risk because the output is a list a person reviews, not an action taken.
Drafting entries from activity, screened against billing guidelines. High value on realization, and the timekeeper reviews every entry before it is submitted.
For in-house teams reviewing outside counsel bills line by line against rules and budgets. Works well because the rules are explicit and the output is a flag, not a decision.
Two engagements from our legal practice. Both produced the clean, consistent operating data that any agent depends on — which is the work that has to come first.
Twopir provided Salesforce customisation and integration services to help us build a robust, compliant, and scalable legal operations platform — connecting case management, document processing, and financial systems into one unified workflow. The result was transformative for how we run case-to-cash operations.
Streamlining case-to-cash operations with Salesforce, AWS and QuickBooks.
Twopir's specialized Salesforce customization enabled efficient integration of third-party systems and streamlined administration and billing, leading to seamless financial operations and enhanced productivity. Automated mass billing and matter management minimized errors across our entire legal workflow.
A 50% efficiency gain from Accounting Seed and Salesforce integration.
Everyone in this market will configure agents for you. The useful question is whether they will tell you when your data cannot support them.
Intake qualification needs very different data from matter summarisation. A single "AI readiness score" is marketing; a per-capability assessment tells you which agents to enable this quarter and which to wait on.
And we have. A withdrawn rollout costs more than a delayed one, and it teaches a firm that AI does not work when what did not work was the data underneath it.
What an agent may see and do is a confidentiality question before it is an AI one. Ethical walls and matter isolation get designed into agent access, not assumed.
Measured baseline, defined success, and a switch-off decision that was made before anyone became attached to the project. "It feels faster" is not a result.
We help growing and mid-market companies solve complex CRM, integration and business system challenges, and we serve enterprise organizations with the same architecture discipline. Firms at that stage need a system that survives the next three years of growth — not one built for the org chart they had last year.
This page covers one service. Each one below goes into the detail a specific team needs — pick the one closest to the question you arrived with.
LitifyAI is the vendor's own set of purpose-built legal capabilities — named agents for intake qualification, conflict checking, matter summaries, total policy, case value and strengths and weaknesses, plus a Damages Assistant, Instant Demands, a Document Assistant, AI time capture and AI invoice review. Litify ACE, its agentic case expert, extends this toward driving work forward rather than only reporting. Agentforce is Salesforce's general agentic layer, available to any org: you define topics, actions and grounding data yourself. Packaged agents are faster where one fits your job; Agentforce is the answer where nothing fits. Most firms end up using both.
Usually not in the same phase, and this is one of the few things we are firm about. During an implementation the data is at its thinnest and the team is at its busiest, so agents configured then read records that do not yet reflect how the firm will actually work. We assess readiness during discovery and normally sequence agents into the phase after go-live, once the case types are settled, stages are being updated reliably and the documents are where they should be. The data work that makes agents viable improves reporting and operations regardless.
That is what a readiness assessment measures, and it is measured per capability rather than as a single score — intake qualification needs very different data from matter summarisation. We look at field completeness on exactly what each agent reads, case type and stage consistency across the record set, document coverage and classification quality, and activity capture rates by role. The output is a viable-now list, a not-yet list, and a specific description of what would have to change to move an agent from the second list to the first.
The org's sharing model, if it is designed to. Agents operate within the access their configuration grants, so this is a confidentiality question before it is an AI question — an agent that can read every matter in the firm is a problem regardless of how well it performs. We design agent access explicitly against the sharing model, including ethical walls and matter isolation, and we enumerate the actions an agent may never take. This is part of the guardrail phase, before any agent is enabled, rather than something checked afterwards.
A human review step catches it, which is why we place one before any consequence and name an accountable reviewer rather than leaving it to whoever is nearest. During a pilot every output is reviewed; at scale that becomes sampling, with monitoring on override rates and output quality over time. Every agent action is logged with what it saw, what it produced and who approved it — which is what you need when a client or a regulator asks how a conclusion was reached. Every agent also has a documented switch-off procedure and a named owner.
Only if a baseline is measured before switch-on, which is why we insist on it. "It feels faster" is not a result and will not survive a partners' meeting. We agree a measurable claim and an exit criterion in advance — time to qualify an intake, time to draft a demand, override rate on suggested time entries — and measure against the baseline at the end of the pilot. If the criterion is not met we switch the agent off and tell you why. That is a successful pilot: it cost weeks rather than a firm-wide rollout.
Sometimes, and the honest answer depends more on your data than on your size. The capabilities with the best value-to-risk ratio for most firms are narrow and checkable: intake qualification, conflict checking, time capture assistance and invoice review. They work because the remit is tight, the output is reviewable in seconds, and the data they need is usually already being captured. Matter summarisation and damages work need much richer records and are where firms most often enable too early. We will tell you which group you are in.
A readiness assessment measures whether your data can support the specific capabilities you are considering, and names the ones that are viable now. If the answer is not yet, you will get that answer with the reasoning.
Readiness per capability · guardrails · scoped pilot · measured against a baseline