AI & Data

When AI stops answering and starts doing

5 min read

An empty meeting room in morning light, its table set beside a corridor running through to another room

For the last few years, most AI at work has waited to be asked. Summarise this document. Draft this email. Find the thing in these notes. Produce a first version of this report.

All useful, and in every case the person is still doing the work. They decide when to start, gather what is needed, ask the question, and take the next step.

That is beginning to change. The newer tools keep working after the first instruction: finding information, using other systems, completing several steps, and coming back when the job is done or when they cannot finish it.

It looks like a change in technology. For a business it is a change in operating model, and the question moves with it. Less how good is the model, and more what is it allowed to do.

Give it a job, not a mandate

“Help with operations” is not a job. Neither is “manage customer enquiries” or “look after reporting”. Where a system is going to act rather than suggest, the boundary has to be drawn a great deal tighter than that.

A job reads like an instruction to a competent new starter: every morning, collect yesterday’s figures from these three systems, flag anything outside these limits, prepare the usual report, and send it to this person for review.

Written that way it can be reasoned about. You know what it has to read, what it is allowed to touch, what finished looks like, and where a person is still required. The narrower the job, the easier all four are to answer.

The hard part is permission

An assistant that writes a paragraph needs almost no access. A system that completes work may need to read a CRM, update a record, create a document, send a message, or start another process. That is a different conversation, and it is the one worth having early.

The rule is ordinary: enough access to do the job and none beyond it. If it only needs to read invoices, it does not need to edit them. If somebody should approve a client email before it goes, the approval belongs in the process rather than in a policy. If it needs three systems, three is the number, even where ten is easier to configure.

None of that is new thinking. What is new is how quickly a permission granted for convenience gets used.

Decide where the handoff is

Human oversight sounds like a safeguard until somebody asks what it means on a Tuesday. Every output? Only the unusual ones? Anything involving money? Anything a client will see?

There is no single answer, but there is a reliable place to look: where judgement or consequence starts. A system can gather the information, prepare the paperwork and mark what looks unusual, and a person decides whether the payment goes. It can assemble the report and draft the commentary, and somebody who knows the client decides what the numbers mean. It can answer from approved material, and hand over the question that falls outside it.

The point is not to keep a person in every click. It is to keep them in the decisions that still deserve one.

Let it stop

One of the more valuable things an automated process can do is decline to finish. People fill gaps without noticing they are doing it. They ask somebody, they make a judgement, they recognise that this one is not standard.

A system may be perfectly capable of proceeding, which is not the same as being right to. A missing document, a figure that does not reconcile, a request that falls outside the normal path: stopping is the correct output in all three, and it has to be designed in. Knowing what success looks like is half the specification. Knowing what not enough information looks like is the other half.

Make the work visible

Once work happens in the background, somebody has to be able to see what happened: what was read, what changed, which rule produced that action, what was sent and to whom, and where the answer came from.

That does not mean an investigation every time a task runs. It means the actions that matter leave enough of a trail to be understood afterwards, which matters most where the process crosses several systems. Automation should remove the person carrying information between them, not the ability to see what moved.

Measure the work that disappeared

The dashboards that come with these tools report tasks run, actions completed and tokens consumed. Those are facts about the system rather than about the business. The useful questions are the familiar ones:

  • How long did this process take before, and how much of that needed a person?
  • How often did somebody have to chase it?
  • How many exceptions and errors came out of it?
  • How much attention does it ask for now?

A system that completes ten thousand actions and saves nobody an hour has not earned anything. A small one that quietly removes five hours of repetitive work a week has.

Start with the boring work

The temptation is to look for the largest job one of these could conceivably do. The better instinct is the other end of the scale: a process that is repetitive, well understood and mildly annoying, where the inputs are known, the next step is predictable, and somebody can tell quickly whether the result is right.

Give it that one. Decide exactly what it may reach, decide where a person is required, make sure what it did can be seen afterwards, and run it long enough to find where the exceptions actually are. Then give it slightly more.

The businesses that get something out of this will not be the ones that delegated the most, fastest. They will be the ones that got good at deciding what to delegate.

Keep reading

More insights.

View all insights

So, what’s slowing your team down?

Let’s talk

Start a conversation

Let’s talk about what’s possible.

Tell us where work is getting stuck or what you want technology to make possible.