How to Decide When an AI Tool Is Worth Keeping
Nielsen Norman Group

How to Decide When an AI Tool Is Worth Keeping

Pressure to adopt AI isn't evidence that a tool helps. The PROVE framework tests one tool against one task and produces a provisional decision you can defend.

Tech workers face constant pressure to adopt AI tools, with little guidance on how to tell whether a specific tool improves a specific type of work. This article introduces PROVE, a lightweight evaluation framework we teach at NN/G, and applies it to a real workflow.

Why It's Difficult to Evaluate AI Tools

Product teams hear from leadership that they need to move faster, use the tools the company has purchased, and keep up with every release, all while spending less money on AI and learning new tools without taking time away from their work. Under that kind of pressure, evaluating a new tool fails in two main ways:

  • We try a tool once, get a plausible result on a demo-friendly use case, and call it productive.
  • We try it on the wrong task, hit a frustrating limitation, and write it off completely.

Neither failure addresses the real question: compared with how you already do this work, does the tool produce an outcome worth adopting?

PROVE stands for Problem Alignment, Risk, Output Quality, Velocity, and Experience. The goal is to evaluate one tool against one task and end with a recommendation you can explain to a manager, teammate, client, or yourself.

PROVE is not meant for organization-wide tool recommendations or for characterizing monetary costs for prolonged use. Instead, PROVE is a framework for a person deciding whether a tool is promising enough to keep using on a specific task. It is designed to alleviate the pressure to rapidly evaluate and adopt new AI tools for work.

Gemini Notebooks Example

To show how the framework works, I'll walk through an evaluation I ran on Google's Gemini Notebooks (formerly known as NotebookLM). Every week, I curate two to four articles worth sharing with colleagues and write a short editorial digest for our team's Slack channel; ideally, it's a paragraph of commentary on each piece, with links. I wanted to know whether Gemini Notebooks could draft that digest for me.

Problem Alignment: Start with the Task, Not the Tool

It's tempting to open a new AI tool and start poking around. Exploration can be fun, but it won't tell you whether the tool is useful. Before you spend money, change a workflow, or ask colleagues to adopt anything, define the task first: what would this tool help you do?

Start with a recurring task that takes meaningful time or effort, then ask whether the tool's capabilities match that task. A tool that looks impressive in a demo may have no useful role in your workflow, while a modest tool may remove a bottleneck you hit every week. AI doesn't exist in a vacuum; it competes against your current process: your current quality, speed, and tolerance for friction.

Gemini Notebooks Evaluation

For my needs, the digest task passed this check easily. It recurs weekly, so any improvement compounds. It doesn't demand publication-quality writing, since the audience is an internal Slack channel. And Gemini Notebooks' core mechanic (upload sources, ask it to synthesize across them) maps directly onto the work: I've already read and selected the articles; I need help turning them into a digest. I also already had access to it through our Google Workspace, so no procurement request was needed.

Risk: Confirm You Can Use the Tool Responsibly

Before spending hours testing a tool, make sure it's appropriate for the data you plan to give it. Is the tool approved by your organization? What kinds of data or information will you put into it? Does the privacy policy permit that use? Are you handling personal information, confidential research data, unreleased strategy, or proprietary code?

When someone is under pressure to adopt AI, creating a personal account or subscribing to an unapproved service can feel like an easy workaround. This practice (often called shadow AI) is widespread: UpGuard's research found that roughly 80% of employees admit to using AI tools their employer hasn't approved. β€˜Widespread’ doesn't mean β€˜safe.’ Using organizational data in an unapproved tool creates risks that can wipe out the time savings you were after.

If you can't use a tool responsibly for the task you have in mind, don't force it into the workflow. Use that finding to redirect the conversation: maybe the team needs a different approved tool, a limited use case with safe inputs, or a formal security review.

Gemini Notebooks Evaluation

The inputs to my AI digest workflow were published, public articles, so there's no participant data or client information. I still checked Gemini Notebooks' privacy documentation, which states that Workspace users' uploads and queries aren't reviewed by humans or used to train models. Low-sensitivity inputs, along with acceptable terms, meant the tool passed. If I were uploading research transcripts instead of public blog posts, this step would have ended the evaluation.

Output Quality: Benchmark Against Your Own Real Work

Once an AI tool meets a need and is safe to use, it's time to evaluate its output against a real benchmark. Pick an actual example of your own work and ask: is the tool's version better than, as good as, or good enough compared with what you already produce?

Quality isn't one universal standard. A rough internal outline can be useful even if it doesn't sound like you; a client-facing recommendation has a higher bar. What matters is

Comments

No comments yet. Start the discussion.