AI vendor demos are optimized to impress; procurement is optimized for feature checklists. Neither surfaces the thing finance actually needs to know: what does a unit of acceptable work cost, and how do you know? These ten questions do. Print them, bring them, and pay attention to which ones produce numbers and which produce adjectives.
Pricing and cost
1. "What does it cost per [our unit of work] — not per token, per seat, or per month?" A vendor who knows their product can convert their pricing to your unit (per invoice, per reconciliation, per email) in one email. If they can't or won't, you'll be doing that math alone after signing — do it before instead. Our pricing guide shows the arithmetic.
2. "Is prompt caching or batch pricing passed through to us?" High-volume finance workflows are exactly where cached-input and batch discounts (often 50–90% on portions of the bill) apply. A vendor building on the major model APIs is getting those rates; ask whether you are.
3. "What happens to our price when the underlying model price drops?" Model prices have fallen consistently. A contract pinned to today's costs quietly converts every future price drop into vendor margin. Negotiate pass-through or a repricing clause.
Accuracy claims
4. "What's your measured pass rate on work like ours — and can we see the logs?" The only acceptable form of an accuracy claim is a measured rate on a defined task with inspectable evidence. "99% accurate" without a denominator, task definition, and logs is marketing. (This is the standard we hold ourselves to: every number we publish ships with raw logs and a one-command verification.)
5. "What counts as a 'pass' in your accuracy number?" Field-level accuracy and document-level accuracy differ enormously: 98% per field on 20-field invoices can mean two-thirds of documents need human correction. Insist on the document/outcome-level number — it's the one your labor costs follow.
6. "Can we run our own held-out test set through it before contracting?" The right answer is an unhesitating yes. Thirty to fifty of your real (anonymized) documents, including your ugly cases, graded against criteria you wrote in advance — the pilot playbook is the procedure. A vendor who resists evaluation is telling you the evaluation's result.
Controls and data
7. "Is our data used for training, and where does it live?" You want, in the contract: excluded from training, defined retention and deletion, known residency. "We take security seriously" is not a data processing term.
8. "What's logged, and can we export it?" Your auditors will ask how you know the control operated. The product should log input, output, model version, and reviewer action per item — exportable, on your retention schedule. See the controls guide for the full control design.
9. "When you upgrade the underlying model, what's your regression process — and do we get notice?" Model swaps change behavior. You want pinned versions, advance notice, and a documented regression suite — the same change management you'd demand for any system touching the ledger.
The closer
10. "Show me a failure." Every AI product fails somewhere. A vendor who can show you a real failure, explain the pattern, and describe how customers route around it has measured their product. A vendor whose product "doesn't really fail" hasn't — or worse, has, and won't say.
Companion download: the AI pilot scorecard — the one-page decision table for scoring vendors against your baseline. For measured cost-per-outcome numbers across models, see the current edition.