satsawat.ai Contact Subscribe

Your Agent Knows You'll Accept. It Doesn't Know If It Helped.

Consumer platforms solved this a decade ago and publish the method. Every proactive-agent benchmark from 2024 to 2026 still scores whether the user said yes.

A person working late in a large dark room, a machine presence beside the desk.

Key Takeaways

Predicting whether someone accepts an interruption is not measuring whether it helped. Netflix, Meta, Pinterest and Spotify all fixed this on their messaging channels years ago, and published how.

The agent field has not followed. Every proactive-agent benchmark I could check, through to ICLR 2026, still scores predicted acceptance against fixed ground truth.

My take: the blocker is not engineering. Acceptance rate usually sits in somebody’s performance review, so changing the metric changes what people are paid.

GitHub’s engineering blog published a piece in November 2025 on teaching Copilot’s next edit suggestions when to stay quiet: train on examples where no edit was right, then tune thresholds to reduce eagerness. That reduces noise. It does not tell you whether the suggestions helped.

A classifier predicting whether you will accept is predicting your behaviour, not measuring whether it made you better off. Usually those move together. The cases that matter are where they come apart.

Four groups, and two of them look identical

Four groups sorted by what the interruption does. Only one is worth paying for, and one does damage.

Some would have acted anyway. Some genuinely needed the nudge. Some were never going to engage. And some are made worse off by being interrupted, the group direct marketing has called sleeping dogs since the late nineties.

You cannot sort individuals into those boxes, because nobody observes both outcomes for the same person. What you can find out is whether that fourth group exists in your system at all.

What it costs when nobody checks

Notification frequency against attrition from the reachable base.

A randomised experiment on 17,500 retail app users, five arms of 3,500, seven weeks (Wohllebe et al., 2021). Two notifications a week lost 6.46 percent of the reachable base against 1.37 for the lightest arm, and the lightest arm finished with the highest engagement of any group.

Three caveats it deserves. The paper does not test that contrast; I do, on their group means. The loss is of users still reachable by push, so switching notifications off counts the same as uninstalling. And the light arm still received two messages, so nobody here got silence.

The platforms already solved this, in public

This is the part that should be uncomfortable. Netflix has withheld a random 2 percent of every non-essential message since 2019 to measure per-message incremental impact. Pinterest published incremental-utility optimisation at KDD 2018, saying plainly they had been wrong to send by click-through. Meta wrote the Instagram objective in do-notation in 2022. Spotify gates in-app messaging on the sign of an estimated per-user effect. LinkedIn shipped a production system last month at +7.20 percent, p = 0.041. Horvitz was pricing the cost of interruption against the value of alerting in 2003.

None of this is new, and I am not claiming the diagnosis. What is new is who has not adopted it.

The agent field is still scoring acceptance

Every proactive-agent benchmark I could check scores predicted acceptance against fixed ground truth. PRISM, at ICLR 2026, gates on a calibrated probability that the user accepts, and contains no mention of randomisation, incrementality or treatment effects at all. A survey of agent evaluation (Yehudai et al., 2025) says outright that causally attributing outcomes to specific decisions is not yet possible.

Two randomised studies suggest what that costs. In METR’s trial, 16 developers worked 246 tasks in repositories they knew well: AI assistance made them 19 percent slower, and they finished believing it had sped them up by 20. In Perry et al., participants with an assistant wrote less secure code and were more confident it was secure. Neither randomised the moment of interruption, only whether the tool was available, so neither is the experiment this asks for. They establish something narrower and worse: when the answer is finally measured, nobody involved could feel it.

And when someone disables your assistant, that write lands in a settings table owned by another team and never joins your decision log. The harm is not unmeasurable. It is unlinked.

Count decisions, not people

Uplift modelling needs many users measured once. Micro-randomised trials need few users measured thousands of times.

The marketing answer is uplift modelling, and at scale it reports averages. One agentic deployment (Jeunen, Hanna and Wheeler) ran across 8.8 million users on a ninety-ten split and published aggregate lifts only, against a control that still received business-as-usual messaging rather than silence.

Health research has a design for the opposite shape, the micro-randomised trial. HeartSteps randomised 37 participants at five decision points a day for six weeks, producing 8,274 raw decision points and 7,540 analysed. That buys power from decisions you already have rather than from more users.

Be precise about what it returns: an average over people and over decision points, moderated by observed context. Not a per-user verdict, so it will not hand you the harmed individuals above. And the effect is defined against the send schedule that produced it, so you cannot measure once and reuse the number under a different policy. In HeartSteps the headline lift came in at p = .06. The unambiguous finding was decay, from a 66 percent lift in week one to undetectable by day 28.

Nobody has applied this design to an agent’s decision to interrupt. That is the gap.

Where I would start

You already have silence you never chose. Threshold changes and rollout boundaries produce it, and reading the discontinuity there gives you a local effect for users near the threshold, who are the ones the model is least sure about. It will not generalise. It costs nothing and it tells you whether the effect is worth arguing about.

Expect the blockers to be organisational. In a regulated firm this reaches model risk before it reaches production, and acceptance rate usually sits in somebody’s performance review, so changing the metric changes compensation. The discontinuity read goes to data governance instead, which is the cheaper door. When you do ask for randomised silence, condition the rate on stakes rather than applying it uniformly.

That is not a way around the ethics. Randomising whether a safety-relevant agent speaks up is a real question, and the health researchers earned these methods inside review boards over fifteen years. The answer is that the alternative is not neutral either: shipping interruptions at high frequency with no measurement of who they harm is also a choice, just an unexamined one.

One honest limit, narrower than I first thought. A field experiment on student financial aid (Athey, Keleher and Spiess) found that estimating treatment effects fully non-parametrically is noisy, and a simple hybrid using a baseline prediction beat it. It did not find that prediction beats causal targeting: nudging the students least likely to file did substantially worse. The criterion is right and the estimator needs help.

And one thing nobody has answered. In a campaign, not sending is a clean control. For an agent, silence changes what the user does next, so the counterfactual is not a quieter version of the same session. It is a different session. Whoever works that out first will have built something the messaging platforms never needed.

Get the next one

If this was worth your time, the next piece will be too. No cadence promised — they arrive when the work is done.

Free, and unsubscribing takes one click. Your address is used to send you the writing — never shared, never sold.