I have told part of this story before.

I built a lot of agents.

Then I audited them and discovered I had built considerably more than I realised.

More than fifty names were scattered across the system. Some were real agents. Some were scripts. Some were good prompts wearing a convincing jacket. Some appeared to exist mainly because I had enjoyed naming them.

I reduced the list to roughly twenty-two and felt as though I had learned the lesson.

I had not.

I thought the embarrassing part was that I had called too many things agents.

The more important part was that every one of them had created a job for me.

Four agents, one briefing

The clearest example was the morning briefing.

I had a daily briefing agent. A weekly one. One that brought different sources together. Another that aggregated reports.

Four agents doing variations of one job.

Each had been sensible when I built it.

I would notice a gap, describe what I wanted and have AI create something to fill it. Because adding a new agent was quick, I rarely stopped to ask whether an existing one should change instead.

That question arrived later, when four versions of roughly the same information started landing in different places.

Now I had to work out which one was current.

Which one had seen the latest information.

Which one I was supposed to read.

And whether the fact they disagreed meant something important or simply meant I had built the same thing four times.

The agents had completed their tasks.

They had handed the actual problem back to me.

Twenty-two agents became thirty-eight jobs

The architecture looked clean on paper.

Multiple agents, each with a role. A hierarchy above them. A chief of staff coordinating the work. Things happening while I slept.

What existed underneath was less elegant.

The twenty-two agents became thirty-eight separate jobs.

Every job had to start somehow. It needed instructions, access to information, somewhere to put the output and a way to behave when something went wrong.

Then the output arrived.

Somebody had to read it.

Somebody had to decide whether it was right.

Somebody had to deal with two agents producing different answers.

Somebody was me.

I had treated the agent as the unit of work.

It was not.

The agent was the beginning of a small supply chain that ended at my desk.

The approval interface

One of the more ambitious agents was built for outreach.

It had 881 contacts queued and ready.

It sent none.

At first, that looked like a technical failure. The agent had done the research, prepared the work and stopped before the useful bit.

But the system was behaving exactly as I had designed it.

I did not trust it enough to send messages without approval.

So it created a queue.

A very impressive queue.

Eight hundred and eighty-one decisions waiting for me.

The system was supposed to reduce my decision load. The biggest thing it produced was an approval interface for me to make more decisions.

That sentence was funny when I wrote it down.

It was less funny when I opened the interface.

Activity is an excellent disguise

From inside the system, everything looked busy.

Jobs were running. Files were changing. Briefings were appearing. Contact records were being prepared. Logs were filling up with evidence that things had happened.

The agents were active.

I was not obviously getting more done.

That distinction is surprisingly easy to miss with AI because the output arrives so quickly.

When a person produces a report, the time it took creates a natural reason to ask whether the report was worth producing.

When an agent produces one in seconds, the cost feels close to zero.

Until there are four reports.

Then forty.

Then a dashboard to organise the reports and another agent to summarise the dashboard.

The cost was not creating the output.

The cost was everything the output expected me to do next.

Agent count was my vanity metric

Every new agent made the system feel more advanced.

It was a workforce. A fleet. A sign that I was getting further ahead.

The count was easy to understand and even easier to be proud of.

It measured almost nothing useful.

An agent that produces work nobody needs is not leverage.

An agent that saves ten minutes but creates twenty minutes of checking is not automation.

An agent that waits for approval on every tiny step has not removed the work. It has changed the shape of the inbox.

The real question was not how many agents I had.

It was what became easier because they existed.

For too many of them, I could not give a convincing answer.

Subtraction was the useful build

So I started removing things.

The four briefing agents became one.

Duplicate jobs were merged. Things that did not need judgement stopped pretending to be intelligent. Agents that existed because of an old system disappeared with the old system.

It felt wrong at first.

Building had always meant adding capability. Deleting something felt like going backwards.

But the smaller system was easier to understand. The outputs had somewhere to go. I knew which briefing to read. I could see which agents genuinely removed work and which ones merely moved it around.

Nobody demonstrates a deletion on stage.

It was still one of the most useful things I built.

The question changed again

I still believe in agents.

I use them. They do things I could not have imagined building for myself a year ago.

But I stopped treating activity as value.

If an agent cannot tell me what human work disappears when it succeeds, it probably has not earned its place.

That gave me a much smaller and more useful system.

It also left me with a harder question.

Once I had fewer outputs, how did I know the ones that remained were true?

Because an agent does not become reliable simply because I deleted the other three.

And a confident answer does not become evidence just because it arrived in the correct inbox.