Agentforce · Data & Knowledge Integration

An agent is only as good as the content it can find.

Agentforce grounding is what turns your knowledge articles, documents and customer records into something an agent can retrieve and answer from — data libraries, search indexes and retrievers, built on Salesforce's data platform. It is the least visible part of an agent programme and the part that sets the ceiling on everything else. Retrieval quality is agent quality.

The Grounding Pipeline
CONTENT SOURCES Salesforce Knowledge Articles · Attachments · Versions Documents & Records PDFs · Files · CRM objects SharePoint · Drive External Systems Web & Public Content TWOPIR GROUNDING LAYER Library & Index Sources · Scope Refresh cadence Retrieval Retriever · Filters Relevance testing Content Quality Dedupe · Rewrite Ownership & review RETRIEVAL TESTED BEFORE TOPICS ARE BUILT 2πr THE ANSWER Sourced Traceable to a real approved passage Current Reflects this week, not last year Permitted Only what this asker is allowed to see AUDIT · INDEX · RETRIEVE · TEST · MAINTAIN
12+
Years CRM & AI delivery
500+
Clients served
250+
Deployments delivered
40+
Consultants & engineers

Trusted by 500+ organizations — including teams whose knowledge base was never written to be read by a machine.

Amberscript
Kacific
Spinify
Zarraffa’s Coffee
Ultra Consultants
RChilli

Grounding Scope

  • Salesforce Partner
  • Data Libraries
  • Retrievers & Indexes
  • Salesforce Knowledge
  • Data Cloud Ingestion
  • Identity Resolution
  • Document Grounding
  • Content Governance
Where Grounding Fails

The agent is not hallucinating. It found your old article

Most “wrong answer” complaints turn out to be retrieval working exactly as designed on content that should not have been there. Grounding problems are content problems.

Three articles answer the same question differently

Retrieval picks one. Which one depends on phrasing, so the agent contradicts itself across conversations and everybody blames the model.

Nothing has been retired since 2019

Superseded articles sit in the index beside their replacements. An agent cannot tell which one is current unless someone has said so.

The content assumes a human reader with context

“Follow the standard process for this region” means something to a five-year veteran and nothing to a retrieval system quoting it to a customer.

The same customer exists four times

Without identity resolution the agent answers from one fragment of a relationship — and confidently tells a ten-year customer they have no history.

Internal-only content is in a customer-facing index

Margin notes, escalation scripts and internal caveats indexed alongside public articles is the grounding mistake with the worst consequences.

Nobody owns keeping it current

A library indexed once at launch decays from the first product change onward. Without a refresh cadence and a named owner, accuracy has a half-life.

What It Means

What grounding an agent actually requires

Grounding is how an Agentforce agent answers from your content instead of from the model's general knowledge. In Agentforce that is built around a data library: you choose the sources — knowledge articles, uploaded files, and other content — and they are indexed on the Salesforce data platform, which creates a search index and a retriever. At conversation time the agent retrieves the most relevant passages and composes an answer from them. It is retrieval-augmented generation, configured rather than coded, and it is what the standard knowledge-answering action runs on.

The thing teams underestimate is that this is a content project wearing a technology name. Creating a library takes an afternoon. Making sure the library contains one authoritative, current, self-contained answer per question — and nothing internal that should never reach a customer — is the actual work. Retrieval is faithful: it will find your contradictory 2019 article and quote it accurately.

Salesforce provides data libraries, indexing and retrievers, Data Cloud ingestion and harmonisation, and the knowledge-answering action. Twopir Consulting provides the audit and the engineering around it: deciding which sources belong in which library and which agents may read them, resolving duplicate customer records so the agent sees one relationship, remediating and rewriting content so it stands alone, testing retrieval against real questions before a single topic is built, and setting the refresh cadence and ownership that keep it true. You get answers that are sourced, current and permission-aware.

We sequence this before the agent build, not after. A retrieval test in week two costs a few days; discovering in user acceptance testing that half your articles cannot answer their own questions costs a launch date. Inside a first Agentforce implementation this is the phase we schedule second, before topics are built. For growing and mid-market companies with a decade of accumulated content, it is usually the largest and most valuable part of an Agentforce programme — and it improves your human agents' lives at the same time.

Source by Source

How each kind of content gets grounded

The third column is the one that decides whether this still works in a year. A source with no named owner and no refresh cadence is a source that will be wrong eventually.

Content sources for Agentforce grounding, how each is made retrievable, and what keeps it current
SourceHow it is groundedWhat keeps it true
Salesforce Knowledge articlesIndexed into a data library with their attachments; retrieval reads published versions.An article owner, a review cycle, and retiring superseded versions rather than leaving them.
Policy & product documentsUploaded or ingested as files, chunked and indexed for passage-level retrieval.A single authoritative location per document, and re-indexing when a version changes.
CRM recordsRead directly by actions, or unified in Data Cloud where profiles are fragmented.Field-level data quality on the specific fields the agent reads — not the whole org.
External systemsIngested into Data Cloud for retrieval, or called live by an action where currency matters more.A decision per source: is a nightly copy good enough, or must this be live?
Internal-only materialA separate library, readable only by employee-facing agents.Library separation enforced by configuration, and re-verified at every release.
Public web contentIndexed where the agent should answer from your published pages.Awareness that the public site is now an agent input, so marketing edits change agent answers.
One question, one authoritative answer, in one library the right agent can read. Everything else is cleanup.
What We Build

Six workstreams that decide what the agent knows

The first two are diagnostic and often change the shape of the whole programme. We would rather find the gap in week two than have a customer find it in month three.

Knowledge Audit & Gap Analysis

We take the questions the agent is meant to answer and check whether your content can answer them — producing a gap list, a contradiction list, and a retirement list before anything is indexed.

  • Question inventory from real requests
  • Coverage mapping question to article
  • Contradiction and duplication detection
  • Staleness and ownership review
  • A prioritised remediation backlog

Retrieval Testing

The deliverable that changes minds. We run real questions against the index and show you, per question, which passage came back — so “our content is fine” becomes a measurement rather than an assumption.

  • Test question set built from transcripts
  • Retrieved-passage review per question
  • Precision scoring against expected source
  • Failure classification: missing, stale or ambiguous
  • Re-run after every content change

Data Library Design & Indexing

Which libraries exist, what goes in each, and which agents may read them. Library boundaries are a security control as much as a relevance one, and they are hard to change once agents depend on them.

  • Library structure and source selection
  • Customer-facing and internal separation
  • Indexing configuration and refresh cadence
  • Retriever setup and filtering
  • Per-topic library permissions

Content Remediation & Rewriting

Rewriting articles so each one stands alone: one question, one current answer, no assumed context, no pointer to a process that lives in somebody's head. Your human agents benefit as much as the AI one.

  • Merging contradictory articles into one source
  • Rewriting for self-contained, quotable answers
  • Removing internal caveats from public content
  • Retiring superseded versions properly
  • Writing the articles the gap analysis found missing

Data Cloud Ingestion & Resolution

Where the agent needs one view of a customer, we ingest from the systems that hold fragments and resolve them into a single profile — so the agent answers about the relationship, not about one record.

  • Source connection and ingestion design
  • Mapping and harmonisation to a common model
  • Identity resolution and match rules
  • Calculated fields the agent needs at conversation time
  • Scope discipline: only what an agent actually reads

Knowledge Operating Model

The part that stops this decaying. Named owners, a review cadence, a route from “the agent could not answer this” into the content backlog, and re-indexing when things change.

What We Connect

Where the content lives, and how it gets in

Grounding is mostly a one-way flow — content in, answers out — and the direction matters, because it decides where a correction has to be made.

Salesforce Knowledge

Published articles and attachments flow one way into the library. Because nothing flows back, a wrong answer is fixed by editing the article — which means the article owner, not the agent team, holds the fix.

Data Cloud

Records from CRM and outside systems are ingested, mapped to a common model and resolved into unified profiles. The agent reads the resolved view; the source systems stay authoritative and unchanged.

SharePoint, Drive & document stores

Policies, manuals and specifications are ingested and chunked for passage-level retrieval. The source system's access rules have to be mirrored deliberately — indexing does not inherit them for you.

ERP, billing & operational systems

A per-source decision: ingest a copy for retrieval, or call live at conversation time. Order status needs live; a product catalogue does not. Getting that wrong is how agents quote yesterday's stock level.

Your public website

Where published pages are indexed, the agent can answer from them — which quietly makes your marketing team upstream of agent behaviour. Worth knowing before a pricing page gets edited on a Friday.

Permissions & the Trust Layer

What the agent may retrieve is bounded by the agent user's access and by which libraries a topic can read. Masking settings decide what reaches the model at all — a grounding decision with a governance owner.

Conversation transcripts

The one genuine feedback loop. Questions the agent could not answer flow back as the content backlog, so the knowledge base improves from real demand rather than from someone's guess about what customers ask.

Retrieval reporting

Which sources actually get cited, and which have never been retrieved once. An article nobody's question ever reaches is either badly written or answering something no one asks — both worth knowing.

How a Grounding Engagement Runs

Test retrieval first, then fix what it proves

Two to six weeks, driven almost entirely by content volume and how much remediation the tests justify. The order is deliberate: measure, then argue about content with evidence in hand.

Step 01 · 3–5 days

Build the Question Set

The questions the agent must answer, taken from real transcripts and tickets rather than imagined — this becomes the standard everything else is measured against.

Step 02 · 3–5 days

Index & Test Retrieval

A first data library over the current content, then every question run against it with the retrieved passage reviewed. This is usually the moment the content conversation changes.

Step 03 · 1–3 weeks

Remediate the Content

Merge contradictions, retire superseded articles, rewrite for self-contained answers, write what is missing — prioritised by what the retrieval test proved was hurting.

Step 04 · 3–7 days

Structure the Libraries

Final library boundaries, internal and customer-facing separation, retriever configuration, and which topics may read which library — verified, not assumed.

Step 05 · 2–3 days

Hand Over the Operating Model

Owners, review cadence, the unanswered-question loop and the re-test procedure, so accuracy is maintained by a process rather than by whoever remembers.

Client Outcomes

Making unstructured content actually usable

Document intelligence and data unification are what grounding is made of. These are Twopir engagements built on exactly those disciplines.

★★★★★
We engaged Twopir Consulting to conduct a Salesforce audit, and their structured, insight-driven approach exceeded our expectations. Their team quickly understood our complex processes, identified critical gaps, and provided clear, actionable recommendations. The audit improved our data accuracy, streamlined workflows, and aligned perfectly with our digital transformation goals.
Kelly Hale Managing Director · Salesforce audit & data accuracy Data Quality
Case Study

Professional Services Firm — AI Document Automation

Automating email attachment processing and Salesforce data routing.

40% Reduction in manual document processing time
2× Admin throughput without added headcount
100% Automated classification & CRM routing
Read Full Case Study
★★★★★
The AI readiness engagement gave us a clear roadmap to operationalize AI across our processes. The team built intelligent ‘Next Best Action’ capabilities using scoring, engagement, and fit models, which significantly improved how we prioritize and interact with prospects. Their understanding of both CRM and AI-driven decisioning made a real difference.
Rubesh J. Consultant · AI readiness & data-driven decisioning AI Ready
Case Study

Professional Events Organization — CRM Integration

Salesforce–MeetMax integration for a corporate networking organization.

100% Elimination of manual data sync between platforms
360° Unified client view across events & CRM
0 Manual reconciliation tasks post-integration
Read Integration Story
Why Twopir

We treat grounding as the actual project

Plenty of partners will configure a data library for you in an afternoon and call the grounding done. The afternoon is not the work; the content is.

We test retrieval before topics are built

A retrieval test in week two either de-risks the whole programme or changes its shape. Either outcome is worth far more than finding out during user acceptance testing.

We will do the content work, not just name it

Most partners hand you a gap list and wish you luck. We merge, rewrite and retire articles as part of delivery, because otherwise the list sits untouched and the agent ships broken.

We scope data work to what the agent reads

An agent programme is not a licence for a two-year data transformation. We resolve the profiles and clean the fields this agent actually touches, and leave the rest for a project that needs it.

Library boundaries are treated as a security control

What goes in a customer-facing index is a governance decision, not a convenience one. We separate internal from external deliberately and re-verify it at every release.

We leave behind an operating model

Owners, cadence and a loop from unanswered questions into the content backlog — because grounding accuracy decays from the day it launches unless somebody owns it.

Common Questions

Grounding questions worth asking early

It is the named set of content an agent is allowed to answer from. You select sources — knowledge articles, uploaded files, and other content — and they are indexed on the Salesforce data platform, which builds a search index and a retriever over them. At conversation time the agent retrieves the most relevant passages and answers from those rather than from the model's general knowledge. Practically, a library is both a relevance decision and a permission boundary: what is in it is what the agent can say.

Only the content in scope for the questions this agent will answer, which is usually a fraction of the library. The retrieval test tells you exactly which articles matter: run the real question set, look at what comes back, and remediate only what is wrong, contradictory or missing. That turns an intimidating “clean up all our knowledge” programme into a specific backlog of perhaps thirty articles. Cleaning everything first is how grounding projects stall before they deliver anything.

Yes — files and documents can be ingested into a data library and indexed for passage-level retrieval, which is how agents answer from policy documents, manuals and specifications. Two practical caveats. First, a long document retrieves as passages, so a policy whose meaning depends on a paragraph three pages earlier will retrieve badly; structure matters. Second, the access rules on the source system do not come across automatically, so who may see what has to be designed into the library boundary deliberately.

Data libraries and retrieval are built on the Salesforce data platform, so the capability is not really separable — the question is how much of the platform your use case needs. An agent answering from indexed articles uses the indexing and retrieval side. An agent that must recognise the same customer across several systems needs ingestion, harmonisation and identity resolution as well, which is a larger scope and a larger cost line. Decide that at the business-case stage; it is the single most common budget surprise in Agentforce programmes.

Separate libraries, enforced by configuration rather than by instruction. A customer-facing agent reads only the external library; an employee-facing agent may read both. Do not rely on telling an agent not to mention something — an instruction is a preference, a library boundary is a control. The other half is content hygiene: internal caveats, margin notes and escalation guidance embedded inside otherwise public articles are the most common leak, and they are only found by reading the articles.

As often as the underlying content changes, which is a business question rather than a technical one. Product documentation that changes with each release needs re-indexing on that cadence; a policy reviewed annually does not. What matters more than frequency is the trigger: someone must know that editing an article changes what the agent says, and that a new product launch means a content task, not just a release note. We set the cadence per source and wire the change trigger into the operating model, then re-run retrieval tests after any significant content change.

Next Step

Give us twenty questions. We will show you what your content actually returns

A grounding review starts with the questions your customers actually ask. We index what you have, run them, and show you passage by passage what an agent would say — before you commit to building one.

Salesforce, CRM & AI delivery for growing and mid-market companies · Contact the team