WritingMarch 18, 2026

How to Choose an AI Tool in 2026: The Evaluation That Ends in a Decision

Two keyboard keys reading A and I, revealed through a torn hole in a sheet of brown kraft paper.

Reviewed by NorwegianSpark Editorial | NorwegianSpark SA

Written with AI assistance and reviewed by the NorwegianSpark SA editorial team.

Last updated: September 2026

Most advice on choosing an AI tool is really advice on shortlisting one. It tells you which categories exist, hands you three names, and stops at the point where the actual difficulty starts. If you want the shortlist, we built one you can work through interactively: the stack matchmaker and the find your AI tool finder both do that job, and our ranked roundup is the reading version.

This page is about the harder half: how to run an evaluation that ends in a decision you can defend, rather than a two-week trial that quietly lapses into a subscription nobody reviews again. That is a procurement problem, not a shopping problem, and the questions are different.

The Failure This Is Designed to Prevent

The characteristic bad outcome is not choosing the wrong tool. It is choosing no tool, slowly. A trial starts with enthusiasm, three people use it in week one, the person who championed it gets busy, nobody ever formally decides, and eleven months later it appears on an expenses review and nobody can say what it did. The money is minor. The cost is that the problem it was meant to solve is still there, and the organisation now believes it "tried AI for that."

Everything below is built to close that loop: a decision, on a date, with a stated reason.

Step 1: Write Down the Job, in the Form of a Sentence You Could Be Wrong About

"We should use AI for customer support" is not a job. "Support agents spend the first four minutes of every ticket finding the customer's previous conversations" is a job, and it is falsifiable — you can go and check whether it is true.

Write the sentence before you look at any product. It does three things at once: it tells you what to measure, it tells you who to ask whether the problem is real, and it protects you from the demo effect, where a tool is impressive at something adjacent to your problem and you buy it anyway. If you cannot write the sentence, the honest conclusion is that you are shopping rather than solving, and the right next step is to watch someone do the work for an hour.

Step 2: Decide Whether This Needs a Product at All

A surprising share of specialist AI tools are a prompt, a file store and an interface around a general model. That is not an insult — the interface is often the valuable part, and building it yourself is rarely free. But it does mean the build-versus-buy question is live in a way it was not for previous software categories, and it is worth ten minutes rather than none.

The question to ask is not "could we build this" but "what is the interface worth to us." If the job is done once a month by one technical person, a general assistant and a saved prompt may genuinely be enough. If it is done fifty times a day by people who should not be prompt-engineering, the product is buying you consistency, permissions, an audit trail and somebody else's on-call rota, and those are worth paying for.

There is a middle path worth knowing about, where you want a tool grounded in your own documents rather than a general model's memory. CustomGPT.ai is built for exactly that shape — a chat agent over a sitemap, a set of PDFs, or a Drive or Notion workspace you already own, answering with citations back to the source page. Whether that is the right answer depends entirely on the sentence you wrote in step one.

Step 3: Ask the Data Questions, and Ask Them of the Documentation

This is the step most often skipped and the only one that can produce a genuinely serious problem. The questions have real answers and the answers are usually written down somewhere the sales conversation will not volunteer.

  • Is your input used to train their models, and can that be switched off? Look for the setting, not the reassurance. If it exists it will be documented.
  • How long is your data retained after you delete it, and after you cancel? "Deleted immediately" and "retained for 30 days in backups" are both defensible; not knowing which is not.
  • Who are the sub-processors? Most vendors publish a list. If a tool runs on a model provider you are not allowed to send data to, the list is where you find out.
  • Is there a data processing agreement, and will they sign yours? For anything touching customer data in a regulated context, this is the gate, and it is a question for your own legal or compliance people rather than for us.
  • Where is the processing done? Relevant if you have residency obligations. Also frequently configurable, and frequently not the default.

Two honest caveats. First, we are not lawyers and nothing here is legal advice — the obligations that apply to you depend on your jurisdiction, your sector and your contracts, and your own counsel is the source. Second, this list is deliberately about questions rather than thresholds, because the thresholds change and a stale number is worse than no number.

The related habit worth building is knowing what is already out there about you and your business before you add another vendor to the pile. Services such as Bitdefender Digital Identity Protection exist to map that footprint.

Step 4: Cost the Whole Thing, Not the Sticker

We are not going to quote prices in this article, because AI pricing changes faster than an article can be maintained and a stale figure is the fastest way to make a bad decision look justified. Go to the vendor's own pricing page on the day you decide.

What does not change is the shape of the cost, and the shape is where the surprises live:

  • The unit. Per seat, per message, per document, per minute of audio, per credit — and credits that expire monthly are a different product from credits that roll over. Two tools with the same headline price can differ several-fold once you divide by the work you actually do.
  • The floor. Minimum seats and annual-only tiers are common at the level where the useful features live.
  • The overage behaviour. Does exceeding the plan cost money, throttle you, or stop the service? All three exist, and only one of them is a billing problem.
  • The setup cost in your own hours. Connecting sources, writing prompts, training people. This is usually the largest number on the list and it never appears on the pricing page.
  • The exit cost. Covered below, because it deserves its own step.

A worked example, using invented round numbers purely to show the arithmetic rather than any real product's rates: a tool at 30 a month that produces 100 usable outputs costs 0.30 an output. A tool at 10 a month capped at 20 outputs costs 0.50 an output, and stops on the twenty-first. The cheaper subscription is the more expensive tool, and it fails on your busiest day. Do this division with the vendor's real current numbers and your own real volume.

Step 5: Design a Trial That Can Fail

A trial that cannot produce a "no" is not a trial, it is an onboarding. Before you start, write down four things:

  • The tasks. Three to five real ones, from your actual backlog, chosen before you see the tool. Include one you already know the right answer to — a known-good case is the only way to calibrate whether the output is right or merely fluent.
  • The people. The people who will actually use it, not the person who found it. Enthusiast results do not generalise.
  • The date. A day on which someone decides. Put it in a calendar.
  • The kill criteria. What result would make this a no. If you cannot name one, you have already decided and the trial is theatre.

Run the same tasks through the shortlisted alternatives, in the same week, with the same people. Comparative evaluation is far more informative than sequential evaluation, because "is this good" is a question nobody can answer and "is this better than that" is a question anyone can.

Step 6: Cost the Exit Before You Commit to the Entrance

The question that separates a reversible decision from an expensive one: if this vendor doubled its price or shut down tomorrow, what would we have to do?

Check, concretely: can you export your data, in a format something else can read, without contacting support? Does the tool hold anything that exists nowhere else — prompts, tuned configurations, a knowledge base, conversation history someone relies on? How deeply is it wired into other systems, and who would have to unpick that?

A tool that produces artefacts you keep is low-risk almost regardless of quality. A tool that becomes the only place a body of knowledge lives is a high-risk dependency even if it is excellent, and it should be judged on a different standard.

Step 7: Set a Review Date and Actually Keep It

Every adopted tool gets a date, six or twelve months out, on which someone asks three questions: is the job from step one still being done, is this still the best way to do it, and is anyone still using it? In a category moving this fast, the answer changes more often than in any software market most organisations have dealt with before.

This is also the step that fixes the failure described at the top. A subscription with a review date cannot quietly become permanent.

Red Flags Worth Walking Away From

  • Accuracy or performance claims with no stated methodology. A percentage without a test set is a marketing number.
  • No published pricing at all for a self-serve product. Sometimes legitimate for enterprise sales; frequently a sign the price is whatever you look like you can pay.
  • A demo that only ever uses the vendor's own example data. Ask to run yours, in the demo, live.
  • Vague answers to the data questions in step three. The answer being complicated is fine. The answer being unavailable is not.
  • Testimonials with no attributable person or organisation behind them.

Where This Framework Is Weakest

Honesty about the method itself: this is overkill for a cheap individual tool. If you are one person deciding whether to pay for a writing assistant, steps two, three and six are largely irrelevant and the whole thing collapses into "try it on three real tasks and see." The framework earns its cost when other people's data, other people's workflows, or a renewal someone else will inherit are involved.

Its second weakness is that it rewards the measurable. Some genuinely valuable tools improve the quality of thinking rather than the throughput of a task, and a trial built around five backlog items will not detect that. Where you suspect this is the case, say so explicitly and decide on judgement — but say it out loud, because "hard to measure" is also the standard excuse for a decision made on enthusiasm.

Disclosure: this article contains affiliate links. If you sign up through them we may earn a commission at no extra cost to you. It does not change what this framework recommends, and no vendor has paid for a mention.

Related Articles

Continue reading

Continue in this collection