The Enterprise Inference Vendor Checklist
On this page
Why a Checklist, Not a Ranking?
Every vendor evaluation watched going wrong made the same mistake early: it started with a benchmark and ended with a surprise. Benchmarks measure what models can do; procurement is about what vendors will do — with your data, your bill, and your exit. Different questions, and only one of them is on a leaderboard.
This is the checklist to bring to any inference vendor call, including a call with us. Five areas, each with the questions to ask and what a good answer sounds like. Bring it to every finalist — the value is in comparing the answers, not in any single vendor’s performance on it.
One note before the list: everything below is verifiable in documents or in a trial. An answer that only exists as a sentence said on a call does not exist. Ask for the page, the policy, or the status history.
Where Does Your Data Go?
The questions: are prompts and completions retained, and for how long? Is any of it used to train models — yours or theirs? Where is data processed, and can a region be pinned? What does deletion look like contractually?
Good sounds like a written policy page answering each of these in plain language, a default of not training on customer data, and a named contact who responds when legal sends the hard version. Vague sounds like “industry-standard practices” and a link to a marketing page.
The platform’s answers live where they should: the privacy policy is the document of record — better to read it than to take this post’s word.
What Does the Bill Actually Look Like?
The questions: are rates published per model, or gated behind a sales call? When savings are claimed, is the figure dated, sourced, and quoted model by model — or one blended average? Can spend be attributed by key, workspace, or team without exporting to a spreadsheet? Is there a committed-use path once the baseline is stable?
Good sounds like a public price list found before anyone contacted you, comparison claims with an as-of date and links to the other side’s pricing, and billing exports matching how the org chart actually looks. The pricing page and the dated comparison snapshot are the platform’s answers to the first two; judge them the same way.
On savings claims specifically: when a vendor quotes one, ask for the range and the conditions, not the hero number. The homepage publishes a forty-five to seventy percent reduction range for adaptive routing, with the mechanism attached — a range with a mechanism is a claim you can test, while a single average is a lobby poster. The invoice teardown shows where the waste that range feeds on usually hides.
What Is the Reliability Story?
The questions: a public status page with real incident history, or a permanently green badge? What happens to an in-flight request when capacity fails — retry internally, or return an error to your code? Rate limits in writing, and what happens at the boundary? Latency numbers measured on your workload during the trial, rather than theirs?
Good sounds like failover described as architecture, with the boring specifics of what a caller sees; status history that includes incidents — a vendor with no visible incidents is a vendor with no visible status page. Everyone has failures; only some vendors let you watch how they handle them.
Reliability is also a cost line, though it never appears on one: every provider error your code absorbs is a retry you pay for. An endpoint failing over internally is quietly cheaper than an identical one returning errors honestly.
What Does the Catalog Look Like — and the Exit?
The questions: how broad is the model catalog, and how fast do new models arrive after release? Is the API OpenAI-compatible, so SDK code survives a switch? And the one nobody asks on the first call: what does leaving look like?
The catalog question matters because model choice is a monthly cost lever that renews monthly — a vendor whose catalog tracks the frontier lets you re-decide without re-platforming. Compatibility matters because it caps switching cost at an afternoon, covered in detail in what “OpenAI-compatible” actually buys you.
The exit question has a sharp edge for fine-tuning: train a model on a vendor’s infrastructure, and do you own the weights, or does the vendor? A fine-tune you cannot take with you is the deepest lock-in in the industry. The platform’s answer is structural rather than rhetorical: fine-tuned weights belong to the customer — which is why the fine-tuning docs spend their effort on training, not on explaining export restrictions.
What Is the Support Reality?
The questions: what channels exist, and what response expectations are written down rather than implied? Who answers — engineers who can read a stack trace, or a tier whose job is to apologise? During the trial, open a real ticket with a real question; the response is the most honest demo available.
Good sounds like response expectations in writing, support staff who can discuss rate limits and model behaviour without escalating every question, and documentation answering the common cases without a ticket at all. The support docs state the channels and expectations; the checklist’s rule applies to the platform too — written down beats said warmly.
Support is where trial behaviour and production behaviour diverge most. Trial answers arriving in minutes? Ask whether that tier continues after signature, and get the answer in the contract.
Which Commercial Terms Matter?
The questions: how long is the commitment, and what happens at renewal — a conversation, or an auto-rollover with a quiet price change? How much notice before rates change, and are existing workloads grandfathered? Do credits expire? On prepayment, what happens to the balance when you leave?
Good sounds like rate changes announced in advance, in writing, current rates always visible on a public page so you can diff them yourself. Bad sounds like a discount contingent on not asking about the list price, and a renewal clause legal finds in section nineteen.
The quietest commercial trap is the expired credit. Generous trial credits with a short fuse create a migration decision under time pressure — precisely when evaluation discipline slips. Treat credits as a way to run the trial, never a reason to shorten it, and note the expiry date on the scorecard next to everything else.
How Do You Score the Answers?
Weight evidence over adjectives — “Enterprise-grade” is an adjective; a dated price list, a status page with scars, and a written retention policy are evidence. Every area above has a document that settles it, and a vendor evaluation is mostly the exercise of collecting those documents and noticing who hesitates.
Then run the trial properly: the same real workload against every finalist, graded on the acceptance checks the product already uses, with cost per completed task as the scoreboard rather than per-token price. The vendor who encourages exactly this test is telling you something — and so is the vendor who steers you toward their benchmark instead.
Keep the scorecard after the decision. The same checklist reviews the incumbent a year later, and incumbents should not be exempt from evidence either.
When Is This Checklist Overkill?
A prototype, a hackathon, a side project — pick anything OpenAI-compatible with a generous free tier and move on. The checklist pays for itself when the decision binds: a contract, a budget, a team whose quarter depends on it.
The middle case — real product, small team, no procurement department — deserves the shortened version: the data-policy page, the public price list, the status-page history. Three documents, ten minutes, and you have already avoided the majority of bad vendor outcomes.
And whatever the scale, keep one habit from the full version: write the answers down. A vendor decision made from memory is a vendor decision made twice.
Related Articles
What "OpenAI-Compatible" Actually Buys You
Every inference provider claims an OpenAI-compatible API. What compatibility actually covers, what it leaves behind, and how to test it in an afternoon.
Where Enterprise AI Spend Actually Goes: An Invoice Teardown
An anatomy of an enterprise inference bill: which line items are legitimate, which are waste, and the questions that find the waste.