Business & Money
Choosing an AI Vendor: The Questions That Actually Matter
By Jim Vernon, Editor, AI Intelligence International · Published 30 January 2026 · Reviewed against our editorial standards · About the author
Every AI vendor demo works, because it was built to. What separates tools that deliver from tools that get quietly cancelled is almost never visible in a demo.
The predictive questions are about your data, your exceptions, model change management, and what happens when the output is wrong.
Key takeaways
- Ask them to run your data, not theirs: Insist on a pilot with a sample of your real cases, including the messy ones.
- Model change management: Ask what happens when the underlying model is updated: do you get notice, can you pin a version, will they re-run your evaluation set, and what is the rollback path?
- Data handling in specifics: Where is data processed, who can access it, is it used for training, how long is it retained, and can you delete it on request with evidence?
- Total cost beyond the licence: Ask what a typical customer of your size spends in year one including implementation, and ask to speak to one.
Ask them to run your data, not theirs
Insist on a pilot with a sample of your real cases, including the messy ones. Vendors will resist, which is itself informative.
Choose the sample yourself and include at least a third edge cases: unusual formats, exceptions, the cases your best staff find difficult. Average-case performance is table stakes; the tail is where value and risk live.
Score the results against a rubric written before the pilot. Otherwise everyone will look at the impressive outputs and ignore the failures.
Model change management
Ask what happens when the underlying model is updated: do you get notice, can you pin a version, will they re-run your evaluation set, and what is the rollback path?
This is the most common cause of post-deployment surprise. A workflow tuned around specific behaviour can degrade overnight with no code change on your side.
A vendor without a clear answer is offering you an unmanaged dependency in a core process.
Data handling in specifics
Where is data processed, who can access it, is it used for training, how long is it retained, and can you delete it on request with evidence? Get the answers in the contract rather than the sales deck.
For regulated sectors add sub-processor lists, breach notification timelines and audit rights. Your risk team will ask; getting it early prevents a three-month delay at the final gate.
Be equally clear internally about what your staff will paste into the tool, because policy is enforced by habit long before it is enforced by controls.
Total cost beyond the licence
Ask what a typical customer of your size spends in year one including implementation, and ask to speak to one. Then assume integration takes longer than quoted.
Watch for usage-based components that scale with success. A tool that becomes expensive precisely when it is working is a budgeting problem you should model now.
Compare against the honest alternative: doing nothing, or doing it with tools you already pay for. Both frequently win and neither is ever in the vendor's comparison table.
Exit and lock-in
Can you export your data, your prompts, your configurations, and your evaluation history in a usable format? Configuration lock-in is the underrated risk, because it is invisible until you try to leave.
Prefer tools that sit alongside your systems of record rather than becoming one. The system of record is the expensive thing to move.
Negotiate a shorter first term with a renewal option rather than a discount for three years. Optionality is worth more than the discount in a market moving this fast.
Deciding between finalists
Score finalists on pilot accuracy against your rubric, integration effort, change management maturity, contract terms and total cost — weighted, and written down before the final demos.
Use the comparison tools to lay the costs out on the same basis; vendors price in deliberately incomparable units and normalising them frequently reverses the apparent ranking.
The five questions that actually filter vendors
Ask where data is processed and stored, whether your inputs train their models, what the retention period is, whether you can export everything on exit, and who is liable when output is wrong. Vendors that answer all five in writing within a week are usually the ones worth piloting.
Two answers should stop a deal for regulated or confidential work: training on customer data without an opt-out, and an inability to name the sub-processors behind the product. Neither is exotic to ask about, and both are common.
Price is the last filter, not the first. A tool that is thirty per cent cheaper and cannot meet your data terms is not cheaper, it is unusable.
Run the pilot so it produces evidence
Define the pilot before it starts: which team, which task, how many weeks, and which two numbers decide the outcome. Pilots without a stopping rule become permanent, and permanent pilots are how shadow spend begins.
Use the same frozen evaluation set across every candidate. Different teams testing different tools on different work produces opinions, not comparisons, and opinions lose to whoever is most enthusiastic.
Write the exit plan during procurement while you still have leverage: export format, notice period, and what happens to your data on termination. Asking for it later is a conversation the vendor has no reason to have.
Frequently asked questions
How long should a pilot run?
Long enough to see real variety, typically four to eight weeks with a defined case volume and a written rubric.
Is a big-name vendor safer?
For data governance and continuity, often yes. For fit to a specific workflow, frequently no. Score both dimensions rather than assuming.
Should we build instead?
Only when the workflow is a genuine differentiator and no vendor addresses it. Building means owning evaluation, maintenance and model change management permanently.
What is the biggest red flag?
Refusal to pilot on your data with your edge cases, followed closely by no answer on model version management.
How long should a pilot run?
Four to six weeks. Shorter misses the novelty drop-off; longer becomes a de facto rollout nobody approved.
Is a single-vendor suite better than best-of-breed?
Suites win on admin, billing and data terms; best-of-breed wins on capability. For small teams the admin saving usually dominates.