Ask the vendor to show you a system already running in production, then ask what breaks it. A vendor who can only show a demo, a slide, or a sandbox has not proven anything you can buy safely. The questions that separate a real vendor from a confident one are boring on purpose: who is liable when the output is wrong, what happens when the model changes under you, how you reconstruct a decision six months later, and whether the person answering has shipped this before or is reading it off a deck. Everything else is atmosphere.
We are an AI vendor. Writing this openly works against us in the short run, because the questions below are exactly the ones that expose a weak vendor, and on a bad day that includes us. That is the point. If you cannot pin us to these answers, do not sign with us either.
The failure rate justifies the paranoia. Gartner predicted at least 30 percent of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, inadequate risk controls, escalating cost, and unclear business value (Source: Gartner press release, 29 July 2024). MIT's NANDA study, drawing on more than 300 enterprise deployments, found that 95 percent of generative AI pilots produced no measurable business return, and concluded the divide was driven by approach rather than model quality or regulation (Source: MIT NANDA, The GenAI Divide: State of AI in Business 2025, reported by Virtualization Review, 19 August 2025). Read together, those numbers say the model is rarely the problem. The vendor and the way they were bought usually are.
The demo is theatre. Ask to see production
A demo is a controlled environment that a vendor has rehearsed. It proves the system can do the task once, on data the vendor picked, with a human ready to intervene off-screen. It proves almost nothing about what happens on a Tuesday in your environment with your data and nobody watching.
So change the request. Ask to see a system this vendor has running in production right now, at another client, at real volume. Ask what its uptime has been for the last ninety days. Ask what the last three incidents were and how long each took to resolve. A vendor with production deployments answers these quickly and specifically, because they live inside those numbers. A vendor without them redirects to the roadmap, the architecture diagram, or the pilot they are "just about to start." That redirection is the answer.
Then ask the harder version: what breaks this system. A practitioner who has run something in production can list the failure modes without flinching, because they have watched each one happen. Vague reassurance that the system is reliable is a signal the person talking has not been on call when it was not.
Make them tell you what they do not do
The fastest way to catch a weak vendor is to ask what their system is bad at, and watch whether they can answer. Every real system has an intended-use envelope: the conditions it was built for and the conditions where it degrades. A vendor who has shipped knows exactly where that boundary sits and will describe it plainly. A vendor who claims the system handles everything is either inexperienced or selling comfort, and both cost you the same in production.
For regulated buyers this matters more than for anyone else. If a vendor cannot state, in writing, what their system is not designed to do, you cannot write an intended-use statement, you cannot scope human review, and you cannot defend the deployment later. In healthcare and pharma, a system used outside the envelope it was validated for becomes your exposure. When a vendor sells the absence of a stated boundary as flexibility, read it for what it is: an unbounded liability with your name on it.
Push hard on liability, and distrust anyone who absorbs all of it
Ask the vendor who is liable when the AI is wrong. The dangerous answer is the reassuring one. A vendor who offers to take on all the liability, including clinical or professional judgment, is either not going to honour it or does not understand the law they are operating under.
Liability decomposes. It does not sit in one place. In clinical AI, for example, clinical judgment stays with the clinician and vicariously with the hospital under Indian law, regardless of what a vendor contract says. No regulatory pathway moves clinical accountability onto a software company, so a vendor promising to absorb it is selling you something they cannot legally deliver. What a serious vendor does take is responsibility for system behaviour: output outside the intended-use envelope, a record attributed to the wrong encounter, retrieval pulling the wrong data, uptime breaches, data protection obligations under DPDP 2023. That belongs in an SLA, with accuracy thresholds and indemnity written to a defined cap, and data breach carved out of the cap. Deployment discipline, running the system inside its envelope with a human review step, stays with you.
The test is whether the vendor can draw those lines cleanly. One who says "we handle all of it" has not thought about it or is hoping you will not. One who decomposes it in front of you is telling you they have been through a real procurement and legal review before.
Ask how you would reconstruct a decision under audit
This is the question that separates a production vendor from a demo vendor faster than any other. Ask: if a regulator, a court, or your own compliance officer asks what the system did on a specific date, can you show them.
That means audit logs, versioned models and prompts, and captured human overrides. Without those, none of the liability allocation above is enforceable, because you cannot prove which failure class you are in. You cannot show the output was inside or outside the envelope. You cannot show whether a human reviewed and signed off. You are asking a compliance officer to trust a system that cannot account for itself, and the correct compliance answer to that is no.
Here is an honest one, applied to us as much as anyone. Observability of this kind, the logs and versioning and override capture, is standard on our deployments. Formal evaluation with expert-labelled ground truth, the harder discipline of measuring accuracy against a validated reference set, does not yet run on every account. It is a current engineering priority, not a finished claim. If a vendor tells you they have full formal eval coverage across every deployment, ask to see the labelled datasets and who labelled them. Most cannot, and the ones who overstate it here will overstate it everywhere.
Reference checks: call the client the vendor did not offer
Every vendor hands you a reference who will say kind things. That call is nearly worthless. The useful call is to a client the vendor did not put forward, ideally one who churned or scaled back. Ask the vendor for the logos of everyone live in your sector, then find your own path to two of them.
When you get a real reference on the phone, skip the satisfaction questions. Ask what broke in the first ninety days. Ask how long support took to respond when it mattered. Ask whether the system that shipped matched the one that was demoed. Ask what they would scope differently if they signed again. A reference who only has praise either had a genuinely clean deployment, which is rare, or is not telling you the parts that would help you.
Structure the contract so a bad fit is survivable
Assume you might be wrong about this vendor, and price that assumption in. A paid pilot with defined success criteria, agreed before it starts, beats an open-ended engagement that quietly becomes a dependency. Write the exit: what you own, how data comes back, what happens to the models and prompts if you leave. A vendor comfortable with a clean exit clause is confident in the work. A vendor who resists it is protecting a lock-in, and lock-in is what turns a mediocre deployment into a multi-year one.
[POSITION NEEDED: Does Nextdot offer a standard paid-pilot structure with pre-agreed success criteria and a defined exit and data-return clause that can be cited as our own practice?]
None of this guarantees a good outcome. It guarantees that a bad one is visible early and cheap to leave. In a market where most projects do not survive the pilot, that is what actually protects you.
Frequently asked questions
What should I ask an AI vendor before buying?
Ask to see a system already running in production at real volume, not a demo: its uptime over the last ninety days, its last three incidents and resolution times, and what breaks it. Ask what the system is not designed to do. Ask how liability is allocated when the output is wrong, and how you would reconstruct a specific decision under audit. Ask for the logos of every client live in your sector so you can run your own reference checks. A vendor who has shipped answers these specifically. A vendor who has not redirects to the roadmap.
How do I know if an AI vendor is bluffing?
Watch what happens when you ask what their system is bad at and what breaks it in production. A practitioner who has run real deployments lists failure modes without hesitation because they have watched each one happen. A bluffing vendor gives vague reassurance, claims the system handles everything, or redirects to architecture and roadmap. The tell is the inability to state an intended-use boundary or to name specific incidents by cause and resolution time.
What are the red flags in an AI vendor pitch?
A vendor who offers to absorb all liability, including clinical or professional judgment they cannot legally take. A system with no stated intended-use envelope. No audit logs, model versioning, or override capture, meaning no way to reconstruct what happened. Only vendor-chosen references. Resistance to a clean exit and data-return clause. Claims of full formal evaluation coverage with no labelled datasets to show. Each one signals either inexperience or lock-in.
How do you verify an AI vendor's claims?
Move every claim from the slide to production. Ask to see the running system, its uptime, and its incident history. Require the liability split in writing, in the SLA, with accuracy thresholds and an indemnity cap. Require observability you can inspect: audit logs, versioned models and prompts, captured human overrides. Run reference checks with clients the vendor did not offer, especially any who churned. Then structure a paid pilot with success criteria agreed up front and a defined exit, so a wrong choice is caught early and cheap to leave.
