Almost every company is under pressure to do something with artificial intelligence. That pressure rarely starts inside the operation: it comes from the board, from a competitor who has announced something, from a conference or from a headline. And because pressure asks for a visible answer, the answer is whatever is visible: a pilot, a licence, a committee, an internal announcement.
Activity is easy to create. Value is a different matter. Gartner expects more than 40 % of agentic AI projects to be cancelled before the end of 2027, because of costs that do not hold up, value that never shows and risk controls that were never there. There is also an MIT report claiming that 95 % of generative AI pilots produce no measurable impact. The figure gets quoted far more often than it gets examined —what counts as failure is arguable— but the order of magnitude matches what we see from the inside.
You do not buy AI. You buy a change of state.
The invoice says licences, hours or platform. What you are actually buying, if you are buying anything, is for your operation to move from one state to another: a task that took four people now takes one reviewing, a decision that took nine days now takes two, a file that came back three times now comes back none.
The test is uncomfortable and it works: if in six months nothing has changed state —same hands, same order, same deadlines, same errors— you have not bought AI. You have bought activity. And activity is billed too.
Agents are workflows with a different name
The word agent has been stretched until it barely means anything. The distinction Anthropic draws in its engineering guide is useful: a workflow follows a fixed path, written by a person, in which the model performs specific steps; an agent decides the path and the tools for itself. Almost everything sold as an agent is the former.
That is not a complaint. Most of the time the former is what you want: it is cheaper, more predictable and easier to audit. But it is worth calling it by its name, because the design, the controls and the cost of a workflow and of an agent are nothing alike. And because a badly designed workflow does not improve by putting the word agent on top of it.
Autonomy with brakes
Delegating to a machine works like delegating to a person: you do not delegate in one block, you delegate step by step. At every step three things have to be decided: what the system may do alone, what it must propose for someone to approve, and what it may not touch at all. Four criteria for deciding:
- Reversibility. A step that can be undone allows more autonomy than one that sends a client an email or moves money.
- Cost of the error. Getting an internal summary wrong is not the same as getting a figure wrong inside a bid.
- Frequency. What happens a hundred times a day needs automatic control; what happens twice a month can take human review.
- Trace. With no record of what was decided and on what information, there is no way to learn from the failure or to defend it afterwards.
This is not a preference of ours. For systems classified as high risk, the European AI Act requires effective human oversight: people able to understand what the system does, interpret its output and stop it. Splitting autonomy step by step is the practical way to comply without freezing the whole process.
Governance that actually runs
Almost every company we talk to has, or is writing, an AI usage policy. Very few have governance that actually runs. The difference is simple: a policy gets signed; governance shows up in the daily work, because somebody authorises, somebody reviews, something gets logged, and there is a written answer for when it fails.
There are serious frameworks to start from: the NIST AI Risk Management Framework and the ISO/IEC 42001 standard, which defines a certifiable AI management system. They are good starting points and poor substitutes for the work: they tell you what has to exist, not how it fits your procurement process, your quality committee or the way you approve a budget. That translation is the work.
Real value, or cost that moved somewhere else
The final question is whether AI has created value or merely shifted the cost and the complexity somewhere else. It happens more than it looks: the time saved by whoever drafts is eaten by whoever reviews; the task leaves the team and returns as system maintenance; the licence is cheap and the total cost of ownership is not.
That is why we measure before starting. With no baseline there is no improvement, only opinion. And we measure both columns: the one that goes down (cycle time, rework, cost per operation) and the one that goes up (review time, incidents, dependence on a supplier). Look only at the first and every project looks good.
The pressure to do something with AI is not going away. The only thing you get to choose is whether you answer it with activity or with a change of state.
Where we would start
If the pressure is already on, the most useful move is not picking a tool. It is picking a process: one that hurts, that repeats and that can be measured. Understanding it as it is today, deciding step by step what gets delegated, building it with the controls in place, and looking at the numbers eight weeks later with the people who use it in the room.
Everything else —the tool, the model, whether it is called an agent or a workflow— decides itself once that part is clear.