All insights
PlaybookAugust 20268 min read

Bulk CRM Enrichment Without the Slop: Cited Values or Nothing

Bulk enrichment is where AI CRM tools either earn trust or destroy it. The cited-or-blank rule, section-by-section research against SEC filings, and why a guessed field is a landmine while a blank one is honest.

By Arvya Team

Visual for Bulk CRM Enrichment Without the Slop: Cited Values or Nothing
Arvya field note · Playbook

Bulk CRM enrichment done right follows one rule: every value written carries a citation to a primary source (an SEC filing, the firm's own website), and any field that can't be sourced stays blank. A blank field is honest; a guessed field is a landmine. This is the moment where AI CRM tools either earn a firm's trust or destroy it, because enrichment is the one workflow where the tool writes hundreds of values nobody will individually fact-check, until one of them shows up wrong in front of a client.

The appeal of bulk enrichment is obvious. At one live deployment, a mid-market advisory firm on DealCloud (~25 bankers), only 7% of sponsor investment-criteria fields were filled in on day zero, and 58% of sponsor records hadn't been touched in over a year. Those are exactly the fields that decide whether a buyer belongs on a list. Filling them by hand is nobody's job; filling them by machine is either a gift or a slow-motion accident, depending entirely on how it's done.

Why the existing options produce slop

Firms have had two ways to bulk-fill a CRM, and both fail in characteristic ways. Data vendors' enrichment feeds go stale: the fund size reflects the fund from two vintages ago, the sector list was scraped once and never revisited, and the feed can't tell you when any given value was last true. You inherit someone else's decay schedule with no visibility into it.

Generic AI enrichment fails worse: it fabricates. Ask a general-purpose model for a sponsor's EBITDA range and it will give you one: fluent, plausible, formatted correctly, and sourced from nothing. The failure is invisible precisely because the output is well-formed. A human reviewing a spreadsheet of 200 confident numbers has no way to spot the eight that were invented. This is what we mean by slop: volume without provenance. The Validity 2025 study (n=602) found 76% of organizations already say less than half their CRM data is accurate and complete; unsourced AI enrichment doesn't fix that number, it launders it.

The cited-or-blank rule

Our answer is a single, non-negotiable constraint: cited values or nothing. If the system cannot ground a value in a checkable source, it writes nothing and says so. This feels like a limitation and is actually the entire product. Consider the asymmetry:

  • A blank field tells the truth. It says “we don't know yet,” which is a research task, cleanly scoped, waiting for a source to appear.
  • A guessed field lies with confidence. It looks identical to a verified one, so it gets used: in a buyer screen, a coverage report, a call with a sponsor who knows their own fund size. One wrong number, caught by the wrong person, ends the evaluation of the tool that wrote it, and fairly so.
  • Citations make review possible at bulk scale. Nobody can fact-check 200 bare numbers. Anyone can spot-check 200 numbers that each link to the filing or page they came from. The citation converts review from re-research into verification.

How section-by-section enrichment works

In practice, enrichment runs section by section rather than record by record. Take sponsor investment criteria: fund size, EBITDA range, equity check size, sector focus. For each sponsor, the system researches those specific fields against primary sources: SEC filings for fund sizes and vintages, the firm's own website for stated criteria and sector language. Each candidate value is staged with its citation attached (the specific filing, the specific page), and a human reviews the batch before anything touches the CRM. Approved values are written and read back from the CRM as receipts, so the enrichment run ends with proof of what landed, not a claim about what was sent.

Sectioning matters because sources cluster by field type. Fund sizes live in filings; sector focus lives on websites; check sizes are often stated nowhere public, in which case the field stays blank and the gap is reported honestly. Running by section means each pass uses the right sources at full depth, instead of skimming everything shallowly per record. The same discipline powers buyer list building, where an uncited buyer simply doesn't make the list.

Opt-in, approval-gated, and measured

Two more constraints keep bulk enrichment from becoming bulk damage. First, it is opt-in: the firm chooses which record sets and which fields to enrich, rather than waking up to a CRM that was “improved” overnight. Second, it is approval-gated like every other write: a human sees the staged values with their citations and approves the batch, or rejects the entries that don't hold up. Arvya never writes to a system of record without a person saying yes.

At that same deployment, this process enriched 40 sponsor records in one week from SEC filings and firm websites, every value cited. That's the honest unit of progress: not “we cleaned your CRM,” but forty specific records that went from mostly blank to sourced and current, with a paper trail behind each field. Intapp's 2024 industry survey found reducing manual data entry is the number-one thing dealmakers want from AI. But wanting the typing to stop is not the same as wanting invented data. The demand is for the work, done to the standard a first-year analyst would be held to.

Measure health over time; never claim “clean”

Finally, resist the word “clean.” A CRM is not a floor you mop once; it is a living record that decays as the world changes, and any vendor claiming to have cleaned it is selling you a snapshot. The right frame is health over time: freeze a baseline (at that firm's day zero, 77% of buyer records were unmatchable and 50,000+ fields were blank), then measure fill rates, citation coverage, and staleness as they move. Progress becomes a chart, not an adjective, and every point on it is backed by verified, cited writes rather than a vendor's assertion.

This is also why the approach stays vendor-neutral: the discipline (cited or blank, approval-gated, read back, measured against a baseline) is the product. The CRM underneath it is your choice, and should remain so. Bulk enrichment without that discipline is just slop delivered faster, and a deal firm's system of record is the last place slop should be allowed to compound.

Keep reading

More from Arvya Insights.

Bring us one live workflow.

See how Arvya reconstructs the work, shows the evidence, and prepares the next action for approval.

Run a live deal