Foxnut Studios

Definition

Buying AI: how to evaluate AI vendors

How to evaluate AI vendors as the operator you are about to become: six product checks before you buy, and the two things buyers should ask for at the contract stage but almost never do.

Reviewed by Ameya Sahasrabudhe and Swati Thakur,

How to evaluate an AI vendor

Evaluating an AI vendor means evaluating a product you will have to run after the sales team has moved on, and the checks that matter are the ones about that afterwards: where the product’s training data comes from, how it has been tested for bias, what the data-privacy terms let the vendor do with your data, how it connects to the systems you already run, how it behaves in production on your work rather than in a demo, and what current customers say when you call them. None of these checks measures how impressive the product is. All of them measure what owning it will be like. Foxnut Studios builds AI systems and hands them over for a living, which in a vendor evaluation puts the studio on the buyer’s side of the table: the questions below are the ones it asks when a client is weighing a purchase, written down so a buyer can ask them alone.

A product you buy is not a person you hire

One distinction sorts most of the confusion in this vocabulary. An AI vendor sells a product: software with model behaviour inside it, priced as a licence or a subscription, run by your team after onboarding ends. A consultant or studio sells judgment and a build: a person or firm you hire, whose output your organisation ends up owning. The two purchases fail differently, so they are vetted differently. A firm is vetted on artefacts, mechanisms and what it refuses to do - the checks for vetting the person or firm you hire are their own list, and almost nothing on it transfers here. A product is vetted on what is inside it, what the contract says about your data, and how it holds up in production. Buyers who apply the consultant test to a product end up judging the salesperson; buyers who apply the product test to a consultant end up asking a two-person studio for a SOC 2 report on software that does not exist yet. Know which purchase you are making before you evaluate anything.

The six checks, and what each one protects

Each check below is something to ask for in writing before a contract is signed. The middle column is the artefact to request; the right column is the failure the check exists to catch. A vendor that cannot produce the artefact is not necessarily hiding something - but the absence is itself the answer, and it belongs in the risk column of the decision rather than in a follow-up call after signature.

The checkWhat to ask forWhat it protects you from
Training data provenanceA written account of what the model was trained or tuned on, and whether your data joins that poolA product whose behaviour nobody can explain, and your records training a competitor’s next version
Bias testingThe tests run before release, on populations like the ones your system will decide aboutDiscovering in production that the product fails systematically for a group of your customers
Data-privacy termsThe clauses on data use, retention and deletion, read by whoever owns privacy internallyTerms that quietly grant the vendor more of your data than the product needs
Integration specsThe documented interfaces, and a named answer to “what does connecting this to our stack take”A purchase that becomes a six-month integration project priced after the fact
Production accuracy and latencyNumbers measured on work like yours, with the conditions stated - not a benchmark from the launch postA demo that was the product’s best day rather than its average one
Reference callsTwo current customers of similar size, chosen by you from a list, not by the vendorBuying on the strength of the one customer for whom everything went well

What buyers ask for, and what they should ask for

The studio’s own record of vendor and build negotiations shows the same pattern on almost every deal: at the contract stage, clients ask for extended support. It is the wrong ask - not because support does not matter, but because it is already built into the commercial agreement, so negotiating hard for it wins nothing that was not already won. What the same clients should ask for, and almost never do, is two things. First, a training and re-training plan for the employees who will actually use the system every day - not a launch workshop, a plan, because the people trained at go-live change roles and leave, and the system’s usefulness walks out with them. Second, the potential expansion of the system against the growth the organisation actually expects - what the product does at twice the volume, in the second market, with the next team onboarded, and what that costs. A vendor’s answer to the expansion question is also the cheapest preview of what being their customer is like in year two.

Cost questions get asked, and one deserves sharper wording than it usually gets. Clients ask for advice on reducing AI token costs, and any competent vendor or builder covers that as a matter of course. The version worth putting to a vendor directly: what in the product’s own design keeps token consumption controlled? Uncontrolled development AI can balloon token costs out of proportion, and a product built that way passes the bill to its buyer. A vendor who can name the mechanism has thought about your bill; one who answers with a pricing page has not.

Evaluating the vendor is not evaluating the output

A neighbouring discipline shares this page’s vocabulary and should not be confused with it. Evaluating LLM output - whether a model’s answers are correct, safe and stable over time - is testing work, done with structured test material against a running system, and it continues for as long as the system runs. This page’s question comes earlier and is about the supplier: whether this company, this contract and this product are the thing to buy at all. The two meet at exactly one point, and it is worth using: a vendor’s response when asked “how would we test this on our own cases before signing” tells you more than most reference calls. A confident product comes with a way to be tried; a fragile one comes with reasons a trial is not representative.

Where an evaluation cannot help

Two limits, stated plainly. First, no vendor evaluation can tell you whether you should be buying at all - that depends on the problem being real, the data being reachable and someone internally owning the outcome, which is the buyer’s own readiness and not the vendor’s quality. Second, the checks above establish what a product is; they do not establish what running it will cost your organisation in attention, retraining and change, which no vendor document states and every operating team discovers. An evaluation done honestly ends in one of three places: buy, with the artefacts filed and the two contract asks made; do not buy, with the failing check named; or not yet, because the evaluation kept surfacing questions about your own side of the table rather than the vendor’s. The third result is the most common one in the studio’s experience, and it is not a failure of the evaluation - it is the evaluation working.

Foxnut Studios works on briefs like this one from Bengaluru and Paris. If you want the shape of that before you talk to anyone, here is how an AI engagement is scoped and priced.