The right test for AI in investment banking is not whether a demo can summarize a CIM. It is whether the product can improve a live workflow without creating a new confidentiality, accuracy, or control problem.
A disciplined evaluation starts with one mandate, one painful loop, and one observable result. It should reveal how the system behaves when evidence conflicts, permissions differ, a banker edits the proposal, or the destination system rejects a write.
1. Start with a recurring operating problem
Choose work the team already performs every week: keeping the buyer tracker current, preparing a weekly client update, coordinating management meetings, reconciling CRM activity, or routing diligence questions. Avoid a broad “AI transformation” pilot. A narrow live process makes adoption, failure modes, and saved effort visible.
2. Demand evidence with every material claim
An answer should show the email, meeting, document, or record that supports it, along with when the source was current. When sources disagree, the system should surface the conflict for review. Confidence without provenance is dangerous on a live deal.
3. Inspect the action boundary
Ask exactly what the product can read, prepare, write, send, or share. Consequential external and irreversible actions should stop for the responsible person. Approval should show the proposed change and its effect, not hide them behind a generic confirmation button.
4. Verify permission and tenant behavior
Test with users who have different deal access. Confirm that the system respects the firm’s actual permission model and does not leak cross-deal context. Document deployment, retention, model, audit, and system-access boundaries rather than relying on broad security slogans.
5. Require a receipt after execution
A successful API response does not prove that the correct CRM field, tracker row, document permission, or calendar object is now in place. The product should read the destination back and preserve the result as a receipt. That is how the team separates a proposal from verified completion.
6. Measure trust, not only output volume
Track how often bankers approve, edit, reject, or ignore proposed work. Review the reasons. A high volume of generated drafts means little if the team does not trust them. The goal is useful work accepted into the live process, with no confidentiality failures and clear evidence when the system abstains.
The pilot should answer one question
At the end of the evaluation, the firm should know whether this workflow became more current, controlled, and easier to run. If the answer is yes, expand to the next adjacent loop. If the answer is unclear, adding more agents or more data sources will not fix the foundation.
Arvya’s recommended starting shape is one live mandate with explicit controls and an agreed security boundary. The product then proves the event-to-receipt loop before scope expands.