Three articles answer the same question differently
Retrieval picks one. Which one depends on phrasing, so the agent contradicts itself across conversations and everybody blames the model.
Agentforce grounding is what turns your knowledge articles, documents and customer records into something an agent can retrieve and answer from — data libraries, search indexes and retrievers, built on Salesforce's data platform. It is the least visible part of an agent programme and the part that sets the ceiling on everything else. Retrieval quality is agent quality.
Trusted by 500+ organizations — including teams whose knowledge base was never written to be read by a machine.
Grounding Scope
Most “wrong answer” complaints turn out to be retrieval working exactly as designed on content that should not have been there. Grounding problems are content problems.
Retrieval picks one. Which one depends on phrasing, so the agent contradicts itself across conversations and everybody blames the model.
Superseded articles sit in the index beside their replacements. An agent cannot tell which one is current unless someone has said so.
“Follow the standard process for this region” means something to a five-year veteran and nothing to a retrieval system quoting it to a customer.
Without identity resolution the agent answers from one fragment of a relationship — and confidently tells a ten-year customer they have no history.
Margin notes, escalation scripts and internal caveats indexed alongside public articles is the grounding mistake with the worst consequences.
A library indexed once at launch decays from the first product change onward. Without a refresh cadence and a named owner, accuracy has a half-life.
Grounding is how an Agentforce agent answers from your content instead of from the model's general knowledge. In Agentforce that is built around a data library: you choose the sources — knowledge articles, uploaded files, and other content — and they are indexed on the Salesforce data platform, which creates a search index and a retriever. At conversation time the agent retrieves the most relevant passages and composes an answer from them. It is retrieval-augmented generation, configured rather than coded, and it is what the standard knowledge-answering action runs on.
The thing teams underestimate is that this is a content project wearing a technology name. Creating a library takes an afternoon. Making sure the library contains one authoritative, current, self-contained answer per question — and nothing internal that should never reach a customer — is the actual work. Retrieval is faithful: it will find your contradictory 2019 article and quote it accurately.
Salesforce provides data libraries, indexing and retrievers, Data Cloud ingestion and harmonisation, and the knowledge-answering action. Twopir Consulting provides the audit and the engineering around it: deciding which sources belong in which library and which agents may read them, resolving duplicate customer records so the agent sees one relationship, remediating and rewriting content so it stands alone, testing retrieval against real questions before a single topic is built, and setting the refresh cadence and ownership that keep it true. You get answers that are sourced, current and permission-aware.
We sequence this before the agent build, not after. A retrieval test in week two costs a few days; discovering in user acceptance testing that half your articles cannot answer their own questions costs a launch date. Inside a first Agentforce implementation this is the phase we schedule second, before topics are built. For growing and mid-market companies with a decade of accumulated content, it is usually the largest and most valuable part of an Agentforce programme — and it improves your human agents' lives at the same time.
The third column is the one that decides whether this still works in a year. A source with no named owner and no refresh cadence is a source that will be wrong eventually.
| Source | How it is grounded | What keeps it true |
|---|---|---|
| Salesforce Knowledge articles | Indexed into a data library with their attachments; retrieval reads published versions. | An article owner, a review cycle, and retiring superseded versions rather than leaving them. |
| Policy & product documents | Uploaded or ingested as files, chunked and indexed for passage-level retrieval. | A single authoritative location per document, and re-indexing when a version changes. |
| CRM records | Read directly by actions, or unified in Data Cloud where profiles are fragmented. | Field-level data quality on the specific fields the agent reads — not the whole org. |
| External systems | Ingested into Data Cloud for retrieval, or called live by an action where currency matters more. | A decision per source: is a nightly copy good enough, or must this be live? |
| Internal-only material | A separate library, readable only by employee-facing agents. | Library separation enforced by configuration, and re-verified at every release. |
| Public web content | Indexed where the agent should answer from your published pages. | Awareness that the public site is now an agent input, so marketing edits change agent answers. |
The first two are diagnostic and often change the shape of the whole programme. We would rather find the gap in week two than have a customer find it in month three.
We take the questions the agent is meant to answer and check whether your content can answer them — producing a gap list, a contradiction list, and a retirement list before anything is indexed.
The deliverable that changes minds. We run real questions against the index and show you, per question, which passage came back — so “our content is fine” becomes a measurement rather than an assumption.
Which libraries exist, what goes in each, and which agents may read them. Library boundaries are a security control as much as a relevance one, and they are hard to change once agents depend on them.
Rewriting articles so each one stands alone: one question, one current answer, no assumed context, no pointer to a process that lives in somebody's head. Your human agents benefit as much as the AI one.
Where the agent needs one view of a customer, we ingest from the systems that hold fragments and resolve them into a single profile — so the agent answers about the relationship, not about one record.
The part that stops this decaying. Named owners, a review cadence, a route from “the agent could not answer this” into the content backlog, and re-indexing when things change.
Grounding is mostly a one-way flow — content in, answers out — and the direction matters, because it decides where a correction has to be made.
Published articles and attachments flow one way into the library. Because nothing flows back, a wrong answer is fixed by editing the article — which means the article owner, not the agent team, holds the fix.
Records from CRM and outside systems are ingested, mapped to a common model and resolved into unified profiles. The agent reads the resolved view; the source systems stay authoritative and unchanged.
Policies, manuals and specifications are ingested and chunked for passage-level retrieval. The source system's access rules have to be mirrored deliberately — indexing does not inherit them for you.
A per-source decision: ingest a copy for retrieval, or call live at conversation time. Order status needs live; a product catalogue does not. Getting that wrong is how agents quote yesterday's stock level.
Where published pages are indexed, the agent can answer from them — which quietly makes your marketing team upstream of agent behaviour. Worth knowing before a pricing page gets edited on a Friday.
What the agent may retrieve is bounded by the agent user's access and by which libraries a topic can read. Masking settings decide what reaches the model at all — a grounding decision with a governance owner.
The one genuine feedback loop. Questions the agent could not answer flow back as the content backlog, so the knowledge base improves from real demand rather than from someone's guess about what customers ask.
Which sources actually get cited, and which have never been retrieved once. An article nobody's question ever reaches is either badly written or answering something no one asks — both worth knowing.
Two to six weeks, driven almost entirely by content volume and how much remediation the tests justify. The order is deliberate: measure, then argue about content with evidence in hand.
The questions the agent must answer, taken from real transcripts and tickets rather than imagined — this becomes the standard everything else is measured against.
A first data library over the current content, then every question run against it with the retrieved passage reviewed. This is usually the moment the content conversation changes.
Merge contradictions, retire superseded articles, rewrite for self-contained answers, write what is missing — prioritised by what the retrieval test proved was hurting.
Final library boundaries, internal and customer-facing separation, retriever configuration, and which topics may read which library — verified, not assumed.
Owners, review cadence, the unanswered-question loop and the re-test procedure, so accuracy is maintained by a process rather than by whoever remembers.
Document intelligence and data unification are what grounding is made of. These are Twopir engagements built on exactly those disciplines.
We engaged Twopir Consulting to conduct a Salesforce audit, and their structured, insight-driven approach exceeded our expectations. Their team quickly understood our complex processes, identified critical gaps, and provided clear, actionable recommendations. The audit improved our data accuracy, streamlined workflows, and aligned perfectly with our digital transformation goals.
Automating email attachment processing and Salesforce data routing.
The AI readiness engagement gave us a clear roadmap to operationalize AI across our processes. The team built intelligent ‘Next Best Action’ capabilities using scoring, engagement, and fit models, which significantly improved how we prioritize and interact with prospects. Their understanding of both CRM and AI-driven decisioning made a real difference.
Salesforce–MeetMax integration for a corporate networking organization.
Plenty of partners will configure a data library for you in an afternoon and call the grounding done. The afternoon is not the work; the content is.
A retrieval test in week two either de-risks the whole programme or changes its shape. Either outcome is worth far more than finding out during user acceptance testing.
Most partners hand you a gap list and wish you luck. We merge, rewrite and retire articles as part of delivery, because otherwise the list sits untouched and the agent ships broken.
An agent programme is not a licence for a two-year data transformation. We resolve the profiles and clean the fields this agent actually touches, and leave the rest for a project that needs it.
What goes in a customer-facing index is a governance decision, not a convenience one. We separate internal from external deliberately and re-verify it at every release.
Owners, cadence and a loop from unanswered questions into the content backlog — because grounding accuracy decays from the day it launches unless somebody owns it.
It is the named set of content an agent is allowed to answer from. You select sources — knowledge articles, uploaded files, and other content — and they are indexed on the Salesforce data platform, which builds a search index and a retriever over them. At conversation time the agent retrieves the most relevant passages and answers from those rather than from the model's general knowledge. Practically, a library is both a relevance decision and a permission boundary: what is in it is what the agent can say.
Only the content in scope for the questions this agent will answer, which is usually a fraction of the library. The retrieval test tells you exactly which articles matter: run the real question set, look at what comes back, and remediate only what is wrong, contradictory or missing. That turns an intimidating “clean up all our knowledge” programme into a specific backlog of perhaps thirty articles. Cleaning everything first is how grounding projects stall before they deliver anything.
Yes — files and documents can be ingested into a data library and indexed for passage-level retrieval, which is how agents answer from policy documents, manuals and specifications. Two practical caveats. First, a long document retrieves as passages, so a policy whose meaning depends on a paragraph three pages earlier will retrieve badly; structure matters. Second, the access rules on the source system do not come across automatically, so who may see what has to be designed into the library boundary deliberately.
Data libraries and retrieval are built on the Salesforce data platform, so the capability is not really separable — the question is how much of the platform your use case needs. An agent answering from indexed articles uses the indexing and retrieval side. An agent that must recognise the same customer across several systems needs ingestion, harmonisation and identity resolution as well, which is a larger scope and a larger cost line. Decide that at the business-case stage; it is the single most common budget surprise in Agentforce programmes.
Separate libraries, enforced by configuration rather than by instruction. A customer-facing agent reads only the external library; an employee-facing agent may read both. Do not rely on telling an agent not to mention something — an instruction is a preference, a library boundary is a control. The other half is content hygiene: internal caveats, margin notes and escalation guidance embedded inside otherwise public articles are the most common leak, and they are only found by reading the articles.
As often as the underlying content changes, which is a business question rather than a technical one. Product documentation that changes with each release needs re-indexing on that cadence; a policy reviewed annually does not. What matters more than frequency is the trigger: someone must know that editing an article changes what the agent says, and that a new product launch means a content task, not just a release note. We set the cadence per source and wire the change trigger into the operating model, then re-run retrieval tests after any significant content change.
A grounding review starts with the questions your customers actually ask. We index what you have, run them, and show you passage by passage what an agent would say — before you commit to building one.
Salesforce, CRM & AI delivery for growing and mid-market companies · Contact the team