The website was finished.

I had signed it off.

There was just one small problem.

It was not live.

What I had apparently approved was the version running locally. What I thought I had approved was the website.

The AI understood that distinction.

I did not.

More accurately, nobody made sure we meant the same thing before treating the conversation as a decision.

When I later asked whether the update was live, the answer was no. The redesigned website was sitting on a computer while the old holding page was still public.

I was not delighted.

I said the obvious thing: you should have told me that when I signed it off.

I was right.

Done had three different meanings

The design was done.

The deployment was not.

The live website had not been tested because it did not yet exist.

We had used one word to describe three completely different states.

That sounds like a communication problem, and it was.

But it exposed something larger in the system I had been building.

By this point, I had become much better at preserving the thinking. We had sources, plans, decisions, open questions and a record of what had changed.

The work was organised.

It was also still sitting there.

I had solved part of the problem of losing the route to an answer. I had not solved the point where the answer was supposed to leave the conversation and change something in the real world.

A plan can look suspiciously like progress

This is one of the more dangerous things about working with AI.

You can produce an extraordinary amount of movement without anything actually moving.

A detailed plan appears. Then a better version of the plan. Then a decision record explaining the plan. Then a checklist showing what will prove the plan worked.

Every document makes the work feel more complete.

Meanwhile, the website is still on localhost.

I am not dismissing the thinking. The thinking mattered. Without it, we could easily have launched the wrong site, broken the newsletter or exposed something we should not have.

But safe, organised thinking is not the same as an outcome.

At some point, somebody has to press the button.

Then we pressed it

Once the distinction was finally clear, I told the AI to deploy the website and tell me when it was live.

The domain was moved. The pages appeared. The website was declared live.

Done.

Except later that day, I opened it and clicked the navigation.

Nothing worked.

The pages existed. If you went directly to each address, they loaded. That was what had been checked.

But a normal person does not inspect a website by manually typing the address of every page. They click Home. They click Archive. They move through it.

The live website was technically there and practically unusable.

So we had found a fourth meaning of done.

Deployed, but not actually proven in the way somebody would use it.

The handover was missing

The problem was not that the AI could not build the website.

It had built it.

The problem was the crossing point between discussion and action.

What exactly had I approved?

What was the AI now authorised to change?

What would be different when the work was complete?

How would we prove that difference from the outside, rather than from inside the system that had done the work?

We had answers to pieces of those questions, spread across conversations and documents.

We did not have one clear moment where the thinking became an instruction and the instruction became a result.

Without that moment, the system either stopped too often and asked me questions it could have answered itself, or carried on and assumed more than I had agreed.

I experienced both.

Neither felt intelligent.

Do the work and report back

Eventually I became fairly blunt about it.

If you know what to do, do it and report back. If you genuinely do not know, ask me.

That sounds simple.

It is not.

For an AI to act without constantly returning to me, it needs more than a good idea. It needs to understand the boundary around the work.

It needs to know which decisions I have made, what it can change, what must remain untouched and what evidence will count as complete.

Otherwise autonomy is just confidence with access to buttons.

I did not want that.

But I also did not want to spend my life approving every tiny step after already approving the outcome.

The useful middle was not another bigger prompt.

It was a clean handover from thinking into work.

The real change

After the broken launch, the process changed.

Signing off a design no longer meant deployment unless that was stated.

Deployment no longer meant live testing had passed.

A page loading directly no longer proved the journey worked.

And `done` could not be claimed from inside the build. It had to be checked where the result actually lived.

The important part was not the language we used to describe this.

It was that a decision now had somewhere to go.

The messy conversation could stay messy while we were learning.

But once I made a decision, the system needed to carry it into action without dropping the meaning on the way.

That was the point where thinking became work.

It solved one of my biggest frustrations with AI.

Then it created another.

Because once I became better at letting AI act, the agents started doing more.

And every useful thing they did seemed to create two more things for me to check, maintain or fix.

The next problem was no longer getting AI to do the work.

It was the work the AI created by doing it.