Over roughly 60 days at a mid-market advisory firm on DealCloud (~25 bankers), a single seat using Arvya daily approved 145 verified CRM updates with a 96% approval rate, automated about 73.5 hours of measured work, built a 54-buyer live list with a cited source on every entry, enriched 40 sponsor records from SEC filings and firm websites in one week, and exported 15 client-ready weekly coverage reports in a single week. Those are the numbers. The more useful output is what the deployment taught us about how trust in an AI system is actually earned inside a deal firm, which turned out to be almost nothing like how it is earned in a demo.
Day zero: what the CRM actually looked like
Before anything ran, we froze a baseline. It is worth publishing because it is the honest starting condition of a real, functioning, revenue-generating advisory firm, not a strawman. On day zero: 77% of buyer records could not be matched to a real firm. More than 50,000 fields sat blank. 58% of sponsor records had not been touched in over a year. Only 7% of sponsor investment-criteria fields (fund size, EBITDA range, check size, the fields that decide whether a buyer belongs on a list) were filled in at all.
None of this made the firm unusual. The Validity 2025 study (n=602) found 76% of organizations say less than half their CRM data is accurate and complete. What made the firm unusual is that they let us measure it, freeze it, and report against it, which is the only way any of the numbers that follow mean anything.
What sixty days of daily use produced
One banker, using the product every working day, approved 145 CRM updates that were written to DealCloud and read back as receipts. The 96% approval rate is the number we watch most closely, because it is a measure of suggestion quality scored by the person with the most to lose from a bad write. Alongside the updates: a 54-buyer live list where every entry carried a citation and 10 were shortlisted; 40 sponsor records enriched in one week, every value sourced; 15 weekly coverage reports exported in one week that went to clients as-is. The ~73.5 hours of automated work is measured from completed, logged work items (briefs generated, updates written, reports exported), not from asking anyone how much time they felt they saved.
Lesson one: trust is earned per-field, not per-demo
The demo won the pilot; it did not win a single approval. What won approvals was the evidence attached to each one. A banker does not decide to trust “the AI.” They decide, one suggestion at a time, whether this fund size with this cited SEC filing behind it is safe to write into this record. Early on, every citation got clicked. A few weeks in, citations on familiar field types got skimmed, and scrutiny moved to novel ones. Approval rates climbed as evidence quality climbed, and the two moved together so tightly that we now treat approval rate as our primary quality metric. The practical implication for anyone evaluating tools in this category: ignore the demo, inspect the evidence attached to ten real suggestions on your own data.
Lesson two: value must arrive before the ask
Every failed CRM initiative in banking shares a shape: it asks for effort now and promises value later. Bankers have seen that movie and they do not stay for the ending. What worked was inverting it. The first thing the banker got each day was a sourced pre-call brief, useful on its own, before any CRM behavior changed at all. The CRM update arrived as a byproduct: the call happened, the notetaker captured it, and the suggested updates showed up already drafted, already evidenced, needing only a yes. Adoption followed the brief, not the pitch. If the first interaction with your tool is a request for the user's labor, you have already lost the median banker.
Lesson three: measure from completed work, not surveys
- Surveys measure enthusiasm, not output. “How much time does this save you per week?” produces a number shaped by mood and politeness. Logged work items (this brief was generated, this update was approved and verified, this report was exported) produce a number shaped by what happened.
- A frozen baseline is non-negotiable. Without the day-zero snapshot, the 145 updates and 40 enriched records would float free of context. Against 50,000+ blank fields, they are honest: real progress, measured, and visibly not the whole job.
- Count only verified units. An update counts when the CRM read-back confirms it landed. A brief counts when it was delivered before the call. Anything the system merely attempted stays out of the tally.
Lesson four: the CRM decays for structural reasons no memo fixes
The day-zero baseline was not the residue of laziness. It was the equilibrium of a system where the people who generate the information (the ones on the calls and in the inbox) are the busiest people in the building, and data entry pays them nothing back. Intapp's 2024 industry survey found that reducing manual data entry is the number-one thing dealmakers want AI to do, which is another way of saying the people closest to the problem have already diagnosed it. Every firm we have met has tried the memo, the Friday-afternoon hygiene push, the analyst assigned to “clean up Salesforce.” The record improves for a month and then the equilibrium reasserts itself, because the structure that produced the decay is untouched.
The only durable fix we have seen is changing what maintaining the record costs: capture the information where it already lives (the call, the inbox), draft the update with evidence attached, and reduce the banker's contribution to an approval. That is the shape of the whole product, and the sixty days were the first sustained test of whether the shape holds under daily, unsupervised, real-deal use. It held. Not perfectly (4% of suggestions were rejected, and each rejection taught us something about evidence quality), but well enough that the firm's record is measurably better than its frozen baseline, with a receipt behind every unit of the improvement. The accumulated, cited record of what the firm knows is what we call the deal brain, and sixty days in, it is no longer hypothetical.