How to evaluate an AI-implementation partner.
Most procurement processes for AI builds optimise for sales presence. Here is the small set of questions and checks that select for shipped work instead.
Most procurement processes for AI builds optimise, without meaning to, for sales presence. The vendor with the slickest deck and the most account managers wins, and the question of whether they have actually shipped a comparable system in production goes unasked. Here is the smaller, sharper set of checks that selects for delivered work instead.
Ask for a system that is in production, not a demo
A demo proves a team can assemble a prototype. Production proves they can handle the unglamorous 80%: evaluation, latency, cost control, failure modes, monitoring, and the long tail of edge cases that only appear at scale. Ask specifically: *Is there a system you built that real users depend on today? Who owns it now? What broke after launch and how did you find out?* The texture of the answer tells you more than any reference call.
Make them show their evaluation
The single clearest signal that a team builds production AI is that they can describe how they measure it. For a RAG system: how is retrieval quality measured, and against what golden set? For an agent: what is the task success rate, and how is regression caught before a deploy? A partner who answers in terms of "it works well" rather than a measurement harness has probably not operated one in production.
Separate the people who sell from the people who build
In services firms the gap between the pitch team and the delivery team is where projects fail. Ask to meet the engineers who would actually staff your build, and ask how the firm protects continuity if a key person leaves. A named, technical delivery lead is worth more than a large but anonymous bench.
Check the claims you are given
Every provider will tell you about their best engagement. The useful move is to verify one claim independently, a case study referenced on the client's own site, a named outcome you can confirm, a talk where they described the work. The point is not suspicion; it is that a claim a third party will corroborate is structurally different from one that lives only in a sales deck. This is exactly the L2 cross-reference bar, and you can apply it yourself.
Insist on a defined scope and a real close
Ambiguous scope is where time and budget disappear. A strong partner will push to define the deliverable, the acceptance criteria, and what "done" means before the work starts, and will treat a dual sign-off at the end as normal, not as friction. A partner who resists pinning down what success looks like is telling you something.
The shortlist test
Run every candidate through five questions: *What did you ship to production, and is it still running? How do you measure quality? Who actually builds it? Which claim can I verify? What does "done" look like, in writing?* The partners who answer all five crisply are a different population from the ones who answer none, and the gap rarely shows up in a sales meeting.
Stay ahead of the AI services market.
One email a month: what's actually being delivered, verified outcomes, rate benchmarks, AI-analysed builds, category shifts. No vendor PR.
By subscribing you agree to our privacy notice. Unsubscribe in one click at any time.
AI builder selection