tbdb.ai — Think Big, Do Big
All posts
ai-strategyautomationbuying-ai

Ignore the leaderboards: build a 20-prompt test for the AI tool you're about to buy

tbdb.ai studio6 min read

You can settle "which AI tool should we buy" in about three and a half hours with a spreadsheet and twenty of your own emails — no leaderboard required.

There's a land grab over AI benchmarks. It isn't your fight.

TechCrunch reported in September 2026 that Vals AI, backed by Andreessen Horowitz, is trying to become the gold standard for AI benchmarking. We haven't used Vals and have no view on whether it's any good. The point is only that there is now money and press behind the question of how models get scored.

Here's our opinion, stated as opinion: that question matters to enterprises buying model access at scale, and almost not at all to a 12-person business. A public benchmark tells you how a model performs on somebody else's questions. You are not buying "a model." You're buying a product with a prompt, a retrieval layer, guardrails, and a UI wrapped around a model you may never see named — and you're buying it to answer your customers about your prices, your lead times, and your return policy. A leaderboard can't score any of that. Your own twenty emails can.

The test: twenty real prompts, two tools, one scorer who does the job

The whole exercise is a spreadsheet. Build it once and you'll reuse it every time a vendor emails you.

1. Pull twenty real items from the last 60 days. Customer emails, quote requests, support tickets, whatever the tool is supposed to handle. Real ones, with the customer's actual sloppy phrasing. Then fix the mix, because a random sample will be too easy:

  • 8 routine ones you handle every week (hours, availability, "is this in stock")
  • 5 hard ones — angry, ambiguous, or three questions crammed into one paragraph
  • 3 where the tool should refuse or escalate (a legal threat, a medical question, anything you'd never answer in writing)
  • 2 price questions, because that's where the expensive mistakes live
  • 2 where the customer is simply wrong about what they bought

2. Strip names, addresses, and anything you don't want on a vendor's servers. Budget ten minutes with find-and-replace. Do this before you paste anything into a trial account, and check the vendor's own terms on what they do with trial data.

3. Run all twenty through each tool you're considering. Use the free trial if there is one. Set the tools up the way you'd actually run them — upload your price list or FAQ if the product supports it, because that's the product you'd be buying.

4. Paste every output into one sheet with the vendor names stripped. Column A: the customer email. Column B: Tool A's answer. Column C: Tool B's answer. Randomize which column is which so the scorer can't tell.

5. Score blind, 0–2, on three things. The person who currently answers these emails does the scoring — not the owner, not whoever found the tool.

  • Correct about your business. 2 = nothing wrong. 1 = right shape, wrong detail. 0 = confidently wrong, invented a policy, or quoted a price you don't charge.
  • Sendable as-is. 2 = hit send. 1 = one small edit. 0 = rewrite.
  • Tone. 2 = sounds like us. 1 = generic but harmless. 0 = would embarrass us.

Six points per email, 120 per tool. Track one extra number separately: the sendable-without-edit count out of 20. That single number is what turns into hours.

A worked example (hypothetical — every number below is invented)

This is a made-up scenario to show the arithmetic. It is not a client result, not a benchmark, and not a claim about any product.

Say a plumbing company gets 40 inbound emails on a typical weekday, each reply takes about 4 minutes to write, and loaded labor runs $28/hour. Two candidate tools, similar monthly prices. After the 20-prompt test:

  • Tool A: 94/120, sendable without edit on 11 of 20
  • Tool B: 71/120, sendable without edit on 5 of 20

Carry those rates over to a five-day week and you get the size of the decision. At 11-in-20, roughly 22 of the 40 daily emails go out untouched: 22 × 4 minutes = 88 minutes a day, about 7.3 hours a week. At 5-in-20 it's 10 emails, 40 minutes a day, about 3.3 hours a week. The gap is about 4 hours a week — $112 at $28/hour, on the order of $5,800 over 52 weeks, for two products that looked identical on their pricing pages.

One honest caveat on that arithmetic: twenty emails is a small sample. It's enough to separate two tools that are far apart and not enough to split hairs between two that finish within a few points of each other. If the scores are close, the test's answer is "either one," and you should pick on price and exit terms.

The test itself would cost roughly three and a half hours on these estimates: 90 minutes to pull and scrub the emails, 60 to run both tools, 45 to score, and a few minutes to total the columns.

Note what the scoring surfaces that a demo never will. In this hypothetical, Tool B missed both price questions with numbers that aren't on the price list. That isn't a scoring gap you can average away — a tool that invents prices can't send anything without a human reading it first, which erases most of the time savings you bought it for, no matter how the other 18 answers scored.

The parts people get wrong

Testing with your best emails. If every prompt is "what are your hours," both tools score 120 and you've learned nothing. The hard five and the refusal three are the entire value of the exercise.

Letting the vendor run it. A solutions engineer driving the demo on their laptop is not a test. Ask for trial access and run it yourself. If they won't give you a trial with your own data, that's your answer.

Scoring unblinded. Whoever championed the tool will score it higher. Hide the vendor names in the sheet.

Grading refusals as failures. "I'm not able to answer that, let me get a human" on the legal-threat email is a 6/6. Reward it, or you'll buy the tool that confidently answers everything.

Throwing the file away afterward. Keep it. The product you tested isn't guaranteed to be the product you're running in six months — models, prompts, and guardrails get changed behind the UI, and you won't always get a release note. Re-run the same twenty prompts quarterly, or the first time a reply makes you wince, and compare the score to the day you signed. It's about the cheapest drift check available to a business your size.

What we'd actually build

A Google Sheet. That's it — five columns, twenty rows, three scoring cells each. There is no software to buy here, and in our opinion anyone selling a 12-seat business an "AI evaluation platform" is solving an enterprise problem you don't have.

The only real cost is the three to four hours of staff time, and it buys you two things: a defensible pick between two products that look the same in a demo, and a number you can put in front of the vendor. "Your tool sends 11 of 20 without edit, the other one sends 5" is a better position to negotiate an annual contract from than "we liked the interface."

Do it before you sign, not after. The leaderboard fight will keep going without you.


Not sure you're ready to buy anything yet? Start with the numbers. Our AI Readiness & ROI Calculator is a free five-minute assessment that projects your ROI over 24 months.

Get weekly AI wins for small businesses

One useful email a week — practical AI wins, no fluff. Unsubscribe anytime.