satsawat.ai Contact Subscribe

Enterprise AI & data advisory

I ship the system and I defend the number

Most enterprise AI advice comes from someone who has never shipped a system, or from someone who has never had to defend a number to a CFO. I have done both for sixteen years — and when I run a benchmark, I publish what it says.

Read the writing Subscribe

Satsawat Natakarnkitkul speaking on stage, holding a microphone.

Writing

Recent writing

All writing

Published, including the inconvenient result

I was shortlisting from the leaderboard, and the ranking did not transfer

I ran thirty-eight forecasting contests — thirteen models, ten datasets, four horizons, minus the two that cannot exist. The model topping the public leaderboard entered all thirty-eight and won none of them, at a mean rank of 9.8 out of 13. Foundation models won thirty overall, so the category claim held; the ranking did not.

I published the result that made my own advice look worse, because a benchmark you only publish when it flatters you is marketing. The method and the code are in the open, and the tool that grades it is on PyPI under MIT.

Read the benchmark agent-report-card on GitHub

New pieces, when there are new pieces

Architecture teardowns, benchmarks run end to end, and honest accounts of what enterprise AI costs and returns. The same writing as on this site, in your inbox instead.

Free, and unsubscribing takes one click. Your address is used to send you the writing — never shared, never sold.

Track record

Sixteen years, and the figures behind them

  • 16+ Years in data & AI
  • $15M+ ARR driven
  • $30M+ AI value realised
  • 40+ Production deployments

The middle two are figures I have defended to the people who sign for them. The last is systems in production rather than pilots, across government, banking, telecom and energy.

Selected work

Systems that went live

Three of these are described on the AWS Public Sector Blog, which outranks a self-reported figure for anyone reading sceptically. Two customers cannot be named, so the sector and the scale are all that is claimed for them.

See all of it

What I am useful for

Four questions that are expensive to get wrong

They are the same four whichever way you buy, and the working for each is already published. Every link below goes to an article, not a case study I wrote about myself.

  • Will it hold at your scale?

    Per-step accuracy multiplies. Chain ten steps at 85% each and end-to-end success is 20% — and the figure in the vendor deck is the per-call one. The work is measuring the chain you actually deployed, then deciding whether to shorten it, checkpoint it, or stop it.

  • Is there anything underneath it?

    Most agent programmes fail on the data, not the model. Three columns named revenue, three different meanings, and an agent picking one with complete confidence — that is not a hallucination problem, it is a vocabulary problem. A catalogue tells you what data exists; it does not tell you what any of it means. The work is deciding what has to be true of your data before an agent is allowed to answer from it.

  • Can you account for it?

    The EU AI Act’s high-risk rules never ask whether your agent is accurate. They ask for named human oversight, six months of logs and incident reporting — an account of what it did. Most deployments were not built to produce one, and a log is architecture rather than storage.

  • Is the number real?

    Business cases get built on figures nobody traced. I have defended $15M+ in ARR and $30M+ in realised AI value to the people who sign for it — and when I ran my own benchmark, the model topping the public leaderboard entered all 38 contests and won none. I published that result rather than the flattering one.

Services

Three ways to work together

Ordered by what they cost you. Most people should start at the top. There is no rate card — I would rather quote against the actual problem than anchor on a figure I invented before hearing it.

  • A workshop

    One or two days with the people who will still be maintaining this after I leave.

    • TimeOne or two days
    • TermFixed. No follow-on
    • You keepA harness your team runs
  • A fixed engagement

    One question, answered on the record, with a start and an end rather than a cadence.

    • TimeThree to six weeks
    • TermOne question, fixed scope
    • You keepA finding, and its method
  • A monthly retainer

    The standing second opinion on a programme that is already running.

    • Time~10 hours a month
    • TermOne quarter, then rolling
    • You keepJudgement, on the record

Full terms, and what I will not take

Boundaries

What I will not take on

Ordered by what each one costs me.

The delivery

I do not bid on the build — not on a retainer, not at the end of a workshop, not as the second phase of a review. Firms win delivery, they are better resourced for it, and an advisor whose next invoice depends on a yes is not giving you advice. Implementation stays with your team or your integrator, and I will tell you when you do not need either. If an engagement ends and the honest answer is that you can do this yourselves now, that is the engagement working.

ASEAN telecom operators

I do not take work with ASEAN telecom operators. It is a sector I have real operating history in, which is exactly what makes it the most expensive line on this list. Everything else is open — banking, insurance, retail, energy, healthcare, manufacturing, government and the public sector.

A decision that has already been made

If the strategy is signed and what is wanted is a name on the cover, I am the wrong person, and you would know it by the second month. I would rather say so in the first call than invoice for three of them.