Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
Software Strategy

AI Vendor Due Diligence Checklist

Last updated:

Why AI vendors need a different kind of checking

The demo went well and the price is within budget. Now comes the part most businesses rush: working out whether this company and this product will still behave well in eighteen months.

Ordinary software due diligence still applies. Where is data stored, who can access it, what happens if the supplier fails. Our list of security questions for software suppliers covers that ground. And before shortlisting, you should already have pushed on accuracy claims with the questions in buying AI software.

AI products add risks traditional software does not have. Many are thin layers over a model from one of a few large providers, so their behaviour can change when that provider updates something. Their output is probabilistic, so a contract promising 'accuracy' means little without a definition. And the market is young, so a meaningful number of today's vendors will be acquired, pivot or close.

This checklist is what we would run on a shortlisted AI vendor. It assumes the decision is worth a few days of effort, which for anything touching customers or core operations it is.

1. The model supply chain

  • Which underlying models does the product use, and from which providers? Is that written down anywhere you can rely on?
  • Can the vendor switch models, and does it notify customers before doing so?
  • When the underlying model changes, does the vendor re-run its evaluation? Will it share results?
  • Is any part of your processing done by a provider in a jurisdiction you need to avoid?
  • What happens to the product if its main model provider raises prices sharply or restricts access?

A vendor that is vague about which models it relies on is either protecting a competitive secret or does not have a grip on its own dependencies. Ask the question in writing and see which it is.

2. Data handling specific to AI

  • Is your data used to train or improve any model, the vendor's or a third party's? Can that be switched off contractually, not just in a settings screen?
  • How long are prompts, documents and outputs retained, including in logs and by sub-processors?
  • Are there sub-processors, including model providers, and are they listed?
  • Can data be processed in a specific region if you need it to be?
  • How is data from different customers separated, particularly in any retrieval or fine-tuning setup?

3. Evidence the product works on your data

Accuracy claims from a vendor's own benchmark tell you very little. What matters is performance on your documents, your tickets, your customers.

  1. Provide a sample of real cases, a few hundred where possible, with known correct answers that the vendor has not seen
  2. Ask them to run the sample and return outputs, not just a score
  3. Score the outputs yourself, including on the hard and unusual cases
  4. Check how the product signals low confidence and whether those signals are trustworthy
  5. Repeat a subset a week later to see whether outputs are stable

The same discipline we apply when evaluating our own systems applies equally to judging someone else's product. If a vendor refuses a test on your data before contract, treat that as the answer.

4. Security, compliance and regulatory position

CheckWhat good evidence looks like
Security attestationA current independent report such as SOC 2 Type II or ISO 27001 certification, with scope covering the product you are buying
Penetration testingA recent summary of findings and remediation, not just a statement that testing happens
Prompt injection and misuseA description of how the product handles malicious input, especially if it reads external documents or emails
Data protectionA data processing agreement and a clear statement of controller and processor roles
EU AI Act positionIf the use could be high-risk, the vendor's view on classification, instructions for use and log access
Incident historyA frank answer on past incidents and how customers were told

5. Commercial and financial health

This is the section businesses feel awkward about and should not. You are about to depend on this company.

  • How is the company funded, and roughly how long is its runway? Private companies will not always say, but the reaction to the question is informative.
  • How many customers of your size and sector does it have, and can you speak to two of them without the vendor on the call?
  • How has pricing changed for existing customers over the past two years?
  • Is pricing tied to usage in a way that could climb sharply if your volume grows or the product's design changes?
  • Who owns the company, and is an acquisition likely to change the product direction?

6. Integration and exit

Exit is the step everyone plans to think about later. Later is usually during a price rise.

  • Is there a documented API, and does it cover the operations you need, not just reading data?
  • Can you export all your data, including configuration, rules, feedback and labelled examples, in a usable format?
  • Who owns any custom prompts, workflows or fine-tuned models created during the engagement?
  • What notice period and transition assistance does the contract provide?

Try the export during the trial rather than taking it on trust. Several of the painful migrations we have handled at SpiderHunts started with an export feature that technically existed and practically produced an unusable file. The contract terms that back this up are covered in our post on AI contract clauses for data rights and exit.

Proportionate diligence, and when to build instead

Not every AI purchase needs all six sections. A writing assistant for a marketing team needs data handling and a quick commercial check. An AI system reading customer contracts or screening applicants needs the lot.

Occasionally the diligence reveals that the product is a thin interface on a general model with your data flowing through several parties, at a price that assumes you will not look closely. In that case a small custom build can be cheaper, more controllable and easier to govern. Our AI integration team will give you an honest comparison, including when the vendor is the better choice, which is often.

Frequently asked questions

What should AI vendor due diligence include?

The usual software checks on security, data location and supplier stability, plus AI-specific checks: which models the product depends on, whether your data trains anything, evidence of accuracy on your own data, how model changes are handled, and a tested data export.

How do we check if an AI vendor uses our data for training?

Ask for it in writing and check the contract and data processing agreement, not just the marketing page or a settings toggle. Also ask whether sub-processors, including the underlying model provider, have any right to use it.

Is a SOC 2 report enough for an AI vendor?

It is a good sign for general security controls, but it does not tell you about model accuracy, prompt injection, training on customer data or model change management. Treat it as one piece of evidence rather than the whole assessment.

How long should AI vendor due diligence take?

For a low-risk productivity tool, a day or two. For a system touching customers, sensitive data or decisions about people, allow two to four weeks, most of it spent testing on your own data.

Keep reading

Shortlisted an AI vendor and want it checked?

We will review the vendor's documentation and test results with you, from the point of view of the people who will have to integrate and support it.

Book a free 30-minute call Get a project estimate WhatsApp us

Related services

What we build for problems like this one

Custom Software DevelopmentDigital Transformation