Tuesday morning, in the weekly status report she is about to send to her leadership, a product manager hesitates over one line: “On this task, I used AI.” The tool gave her nothing in particular that week. But her company tracks an adoption metric, and that line will be counted whether she writes it or not. So she writes it.
Thousands of product teams replay this scene every week. A screenshot from a chat assistant slipped into a document, a tool name dropped into a status report, not because the work was better for it, but because you have to prove you’re keeping up. This reflex has a name: performative AI usage. It thrives because almost no one, inside the company, actually knows how to evaluate an AI tool on real work.

Measuring adoption before measuring value
That is the diagnosis Caleb Sponheim makes. Sponheim is a researcher at the Nielsen Norman Group (NN/g), the digital usability firm founded by usability pioneers Jakob Nielsen and Don Norman. In an article published on August 7, 2026, he describes a drift that has become ordinary: teams evaluated on their AI usage, with no clear definition of what successful usage even looks like. The result, he writes, is that companies “measure adoption before they’ve measured value.”
The line sums up the trap many organizations have fallen into over the past two years. Told to show a usage rate for generative AI (systems capable of producing text, code, or images from a plain-language instruction, a prompt), teams try tools with no grid for judging whether they actually add anything. Sponheim identifies two mirror-image failures: testing a tool once, on a flattering case picked for a demo, and declaring it productive; or testing it on the wrong task, hitting its limits, and dropping it altogether. Neither approach answers the only question that matters: compared with how the work gets done today, does this tool produce a result good enough to justify adopting it?
The phenomenon has a second, quieter face. According to a study by the firm UpGuard, relayed by NN/g, roughly 80% of employees admit to using AI tools their employer has not approved. This shadow AI (the AI-era equivalent of shadow IT, which corporate IT departments have tracked for twenty years: digital tools used outside official channels) thrives precisely because no one has supplied a method for judging a tool before using it.
Five criteria for evaluating an AI tool on a single task
Faced with that gap, Sponheim proposes PROVE, the acronym for the five criteria he formalized: Problem, Risk, Output, Velocity, Experience. The framework does not claim to settle, at the scale of an entire organization, which AI tool to roll out. It aims for something more modest: a defensible decision about one specific tool, applied to one specific task. Spelled out, the five criteria read as follows:
- Problem: start from the recurring task that costs real time, before picking a tool. AI is not evaluated in the abstract; it competes against the current process, its quality, its speed, the friction you already tolerate.
- Risk: confirm, before any test, that the usage is approved and that the data you feed the tool raises no confidentiality issue. The fact that 80% of your colleagues use a tool says nothing about its safety.
- Output: compare the tool’s output against your own real work, not against an idealized example. The quality bar shifts with the stakes of the task.
- Velocity: time the total duration, from launch to publication, corrections included. A three-second generation can still open up a thirty-minute editing job.
- Experience: separate one-time friction tied to learning from friction that recurs with every use. A good tool fits the work naturally enough that you keep using it once the novelty wears off.
The case of a weekly digest, scored criterion by criterion
Sponheim illustrates his method with a concrete case: writing a research digest published every week on Slack, which used to take him twenty-five minutes. On problem and risk, the tool tested, Gemini Notebooks, scores full marks: a repetitive task, modest publishing standards, public sources with no confidential data, a tool already available inside his work environment. On output, though, the score caps at four out of five: the text is accurate and its sources correctly cited, but it “reads like a report” rather than a personal voice, and needs a rewrite before publishing. Total time drops from twenty-five to about ten minutes, a net gain comparable to the ones we described in what the best product managers actually gain from using AI, but the workflow expands from three steps to six: read, upload, write the prompt, copy, reformat, publish. Two extra handoffs every time.
The final call is neither a rejection nor a permanent adoption: a one-month trial, with a caveat spelled out in black and white. The PROVE framework does not produce a binary verdict but a four-part summary, which Sponheim recommends keeping on record: what was evaluated, what was observed, what you decide to do, and the main caveat that remains.
What this framework shifts, beyond the tool
PROVE’s real interest has nothing to do with its acronym, or even its novelty: published only days ago, it has not yet been tested at scale. What it shifts is the question organizations ask about AI. For two years, most have measured adoption rates, activated licenses, mentions in status reports: compliance indicators, not value. When a compliance indicator becomes a criterion for individual evaluation, teams quickly learn to satisfy it without changing their work in any real way. That is the performative usage described above, one instance of Goodhart’s law, well known to economists: an indicator stops being reliable the moment it becomes, in itself, a target to hit.
So it is the burden of proof that shifts. The task is no longer to justify, after the fact, a usage already decided by leadership, but to document, before rolling a tool out widely, what it actually changes about a real task. That is also what consultant John Cutler argued in our pages about trust in AI projects: without a clear social contract between leadership and teams, adoption fails no matter how refined the tool. PROVE supplies one building block for that contract: a shared language for saying, evidence in hand, why you keep a tool, why you give it another month, or why you drop it.
The limits of a framework barely a week old
Sponheim flags the limits of his own method. PROVE is not built to settle a decision at the scale of an entire organization, nor to price out the cost of a subscription over time: it compares one tool to one task, nothing more. The scores it produces, he insists, “reveal a pattern, they don’t replace judgment.” A single run on a single set of inputs remains “a screen, not a verdict.”
One caveat the article does not state outright, but that use will quickly reveal: a framework built to resist pressure to adopt can, by a kind of pendulum effect, feed a reflexive distrust, the kind that leads teams to pile up criteria in order to never change anything. Sponheim partly anticipates this by noting the method should be revisited whenever a task changes, but discipline alone does not always contain a culture already braced against change. And because the method has existed publicly for less than a week as these lines are written, no large-scale field feedback yet exists to confirm it delivers beyond the single case its author chose to document.
What’s left to do, starting Monday
Nothing stops a team from picking this up in the meantime, at its own scale. The next time a generative AI tool shows up inside a workflow, the question is not “should we adopt it?” but “on which specific task, compared with what, at what total time cost, and with what friction?” Writing the answer in four lines is already enough to turn pressure to adopt into a decision you can defend to your team, and reread in six months.
Tuesday morning’s product manager may not need to delete her line after all. She needs a different sentence behind it: not “I used AI,” but what she actually measured while doing it.
Sources
- How to Decide When an AI Tool Is Worth Keeping · Nielsen Norman Group, Caleb Sponheim, August 7, 2026
- Why Trust Decides the Fate of Your AI Projects · Impact Factories
- What the Best Product Managers Actually Gain From Using AI · Impact Factories