AI Hype vs AI Value in Vendor Pitches
Last updated:
The pitch you have probably sat through
The slides are good. There is a demo where an agent reads an email, checks three systems and drafts a perfect reply in eight seconds. There is a number: '80% reduction in handling time'. There is a logo wall. And there is a price that seems reasonable until you notice it is per seat, per month, for three years.
Most AI vendors are not dishonest. But the market rewards confident claims, and the gap between what a product does in a demo and what it does with your data on a Tuesday is often wide. Your job in the meeting is to measure that gap.
Phrases that usually mean less than they sound
| What the pitch says | What it often means | What to ask |
|---|---|---|
| 'Up to 80% time saving' | The best case from one customer | What was the median, and against what baseline? |
| 'Autonomous agent' | Automated steps with a model deciding some branches | Which actions does it take without a human approving? |
| 'Learns your business' | Retrieval over documents you upload | Does it change behaviour over time, and how is that tested? |
| 'Enterprise-grade accuracy' | No specific number | Accuracy on which task, measured how, on whose data? |
| 'Proprietary AI' | A well-known hosted model with a prompt | Which model family is underneath, and can you switch? |
| 'No integration needed' | Manual export and import | How does data get into and out of our systems? |
None of these answers is necessarily bad. A product built on a hosted model with good retrieval can be excellent. The problem is only when the language suggests something the product does not do.
What real value looks like
- A baseline. 'Customers previously took 11 minutes per ticket; after three months the median was 6.' That is a claim with a shape.
- An honest failure description. A vendor who can tell you where their system goes wrong, and how often, has measured it. One who cannot, has not.
- A pilot on your data. Not the demo set. A few hundred of your real documents, emails or records, scored against what your team actually did.
- References you choose. Ask for a customer of similar size in a similar sector, and ask them what did not work.
- Pricing tied to something measurable. Outcome-based pricing, per resolved ticket or per processed document, is becoming more common in 2026. It aligns incentives, though check how 'resolved' is defined.
Test it yourself before signing
The cheapest protection is a structured trial. It does not need to be elaborate.
- Pull 200 to 300 real, recent examples of the task, including awkward ones
- Record what a competent member of staff did with each, as the answer key
- Run them through the product with no special preparation by the vendor
- Score the output: correct, correct with minor edits, wrong, wrong and harmful
- Time how long the human review of the product's output takes
The 'wrong and harmful' column is the one that decides it. A tool that is right 90% of the time and harmlessly wrong the rest is a bargain. One that is right 95% of the time and confidently wrong in ways staff will miss is a liability. Our guide to evaluating language models for business explains the scoring in more detail.
Commercial terms that reveal confidence
How a vendor structures the contract tells you what they believe about their own product. A vendor confident in value will usually accept a paid pilot with a clear exit, monthly terms after it, and a success measure written into the agreement.
Be wary of long minimum terms before a pilot, per-seat pricing for a tool whose value is per task, and vague data terms. On data, ask directly: is our data used to train your models, where is it processed, and what happens to it when we leave? We cover the lock-in side of this in AI integration vendor lock-in.
Regulation is now part of the pitch
With EU AI Act obligations phasing in, some vendors now lead with compliance. That is useful if it is specific. Ask which risk category they believe your use falls into, what documentation they provide, and what they expect you, as the deployer, to do yourself. A general assurance that the product is 'compliant' does not transfer your responsibilities to them.
When hype is not the problem
Sometimes the product is genuinely good and still the wrong purchase. If your process is unclear, your data sits in spreadsheets or the task only happens twenty times a week, even an excellent AI tool will disappoint. The questions in what to ask AI software vendors help with that side of the decision.
When SpiderHunts reviews an AI proposal for a client, the first question is not about the vendor at all. It is whether the underlying task is well defined and frequent enough to be worth automating, by anyone. If you are weighing a vendor against a custom AI integration, the same trial applies to both of us, and we are happy to be tested that way.
Frequently asked questions
How can you tell if an AI product is just a wrapper?
What questions should I ask an AI vendor?
Is outcome-based pricing for AI a good idea?
How long should an AI product trial last?
Sitting on an AI proposal you are unsure about?
Send it over. We will give you a plain reading of what is solid, what is marketing and what to ask before you sign, with no obligation to use us for anything.
Related services
What we build for problems like this one