satsawat.ai Contact Subscribe

Whoever Picks the Metric Picks the Product

The highest-leverage decision on an AI product is which number gets to say it works. It is usually chosen by default, often wrong, and after the demo you cannot take it back. A worked example where the leaderboard's second-best model is the one to ship.

A lone figure stands in a dark hall between a glowing blue path and a warm gold path that split apart on the floor

Key Takeaways

The metric that says “it works” is the real product decision. It gets chosen by default, before anyone treats it as a decision, and the default is often the wrong one.

The leaderboard’s top model is often the wrong one to ship. Score models by the decision they drive at your real cost, not by average accuracy, and the ranking can flip.

You cannot win the argument after the demo exists. Write down the one number a launch has to clear, at your real operating point, and get it agreed before a line of product code ships.

My take: As building gets cheap, the durable job in data science is not building the model. It is owning the definition of “does it work” and locking it as a release gate before the build. Claim that number, or inherit the one nobody chose.

The Decision Nobody Makes on Purpose

Here is a decision that gets made without anyone deciding it. A team wants their retrieval system to do better, so they widen the search and watch recall climb. Recall is the share of times the right answer is somewhere in what the system pulled back, and a wider search can only raise it. But ship that wider search and the product answers no more questions correctly than before, and sometimes fewer, because more results can crowd out the right one. Finding the answer and using it are two different things. Someone chose recall as the number that counts, by default, and that choice had already decided the product would not improve.

The biggest choice on an AI product is which number is allowed to say it works. In 2026 almost no one makes it on purpose. Models now do more of the building, but judging what to measure is the part they cannot do, and whoever owns it decides what ships.

From Training to Checking

One shift explains it: the move from traditional machine learning to agentic AI, software that calls a model to plan and act over many steps. It changed where the difficulty sits. In classic machine learning you trained the model, and training was the expensive part. Checking it had a standard playbook: hold out a test set, read the accuracy, run an A/B test, and you knew whether it worked.

Classic machine learningAgentic AI
The hard partBuilding the modelChecking whether it worked
How you checkHold out a test set, read the accuracy, run an A/B testDefine a good run, sample real traces, judge by hand, and redo it after every model update
Where the value sitsIn buildingIn the judgment

Agentic AI turns that around. You call a model through an interface, so the building is cheap. But the system runs many steps and behaves differently on every run. There is no clean, stable accuracy number the way there was. You define what a good run looks like, sample real traces, and judge them by hand. Each model update makes you do it again. Now the checking is the hard part.

That one shift did two things. It broke the old do-everything role into new lanes, and it made a launch turn on one question: which number gets to say the agent worked.

The Job Is Choosing What Good Means

So what is the job now? Choosing what “good” means for this product and turning it into a number you can measure. Then running the experiment or a causal check when you cannot randomize, and reading the result honestly enough to act on it. Naming that number is the step that decides the most and gets noticed the least. Building the harness that runs the number is becoming an engineering job. Deciding what it should measure is yours.

None of that got cheap. In one careful long-analysis benchmark, the best model still scores under half, and it drifts further off the longer the task runs. And the research mapping what AI automates in data science puts the framing, the reading, and the deciding on the human side.

The job titles say the same thing. PlayStation’s data scientist role is literally called Measurement, Experimentation and Causal Inference. The postings ask you to frame ambiguous questions as measurable ones and to report so the work leads to a decision instead of a dashboard update.

What the Deciding Work Takes

Four questions decide whether a number means anything:

  • What does “good” mean for this business, in something you can measure. Reaching the right document is not the same as answering the question, as the recall trap shows.
  • Did the change cause the result, or did something else move at the same time. Without that, you are guessing.
  • Is the test set the real world or the convenient one. A sample built from what was easy to collect will point at the wrong answer and make it look safe.
  • How sure are you. A single number with no range around it is a guess.

None of this is exotic. It is ordinary data science, the measurement craft that was always the job. What changed is the thing being measured. The data is now traces from a system whose behavior moves run to run and release to release, so the measurement is harder and it goes stale sooner.

The Metric Decides What You Ship

Every team faces the clearest version of this: which model to ship. That looks like a technical call, but the metric you rank the models on is the product decision, made early and usually by accident.

A public benchmark ranks models by average accuracy. But your product may only act when the model is confident, and hand the rest to a person. Say two models cost the same per call. Here is what happens across 100 calls:

ModelLeaderboardActs onRightWrongPassed to review
Model A#1, 82%almost every call82180
Model B#2, 78%60 calls58240

Model A is overconfident and acts even when it is likely wrong. Model B knows when it is unsure and holds back. If a wrong action loses a customer and a passed call costs a few minutes of review, Model B is the one to ship, and the leaderboard ranked it second. Change the cost of a wrong action and the answer changes with it. Those numbers are illustrative, but the reversal is common. Score the models the standard way, then score them again by the decision they drive at your real budget, and the two rankings often disagree. What you ship depends on the metric you grade with, more than on the models themselves. Picking that metric is itself a measurement decision, usually made without thinking.

It shows up again after the demo. In one survey, about 95 percent of enterprise AI pilots moved no profit number, though the sample is small. The report calls that a learning problem, a sign teams have not worked out the process yet. The likelier story is a measurement gap: no one owned the job of saying whether the deployed system changed a number the business cares about, and no one owned the decision of what to do if it did not.

Own the Number Before the Build

If you lead a data science team, stop competing with the AI engineers on building. Put your team on the number that decides whether their work ships, and own it before the build starts.

You own the definition of “does it work” for each product, in numbers the business recognizes. You decide whether the test is sound and whether the sample looks like reality. The uncertainty on every reported number is yours, and so is the call that follows it: ship, hold, or accept a known error rate at a known cost. And you make that work visible to people who would rather ship the demo and find out later.

To take it rather than only claim it, turn the definition of “does it work” into a release gate. Write down the one number a launch has to clear, the decision metric at your real operating point, and get it agreed before a line of product code ships. Once the demo exists, that definition gets argued against something people already want to launch, and you will lose, so settle it while changing your mind is still cheap.

The experimentation literature has a name for the discipline, the overall evaluation criterion: a small set of metrics chosen in advance, with guardrails that hold a launch when it breaks something that matters. The teams that run it well, Netflix among them, put nearly every change through a test where the measured effect decides whether it ships, and the gate is allowed to refuse. Set that threshold before the model is scored. A team that waits until it sees the results is grading to confirm what it already wants.

Be honest about the gap between this and most data science jobs today. Plenty of the work is still dashboards and pipelines, and plenty of people who frame and measure well still get overruled when the call is made. You earn the deciding work by being right in public, again and again, until the team waits for your judgment before it decides. The shift is real, but no one hands you the deciding work.

Pick the Number First

Every AI product ships on a number. Someone chooses that number, and if no one chooses it on purpose, it gets chosen by default and often chosen wrong. The next model will do more of the building every year. What it cannot do is decide what a good result means for your business and stand behind the call when the number is not clean. That decision is made early, at the whiteboard, and it does not come back after the demo. Pick the number before the build, or inherit the one nobody chose. Building got cheap. The metric still decides.


Read more: How Netflix runs the metric as the launch gate  ·  The overall evaluation criterion

Get the next one

If this was worth your time, the next piece will be too. No cadence promised — they arrive when the work is done.

Free, and unsubscribing takes one click. Your address is used to send you the writing — never shared, never sold.