The first useful piece of evidence was a 404.

One of my agents gave me a database ID for a record it said existed.

The answer looked completely normal. The ID had the right shape. The explanation made sense. There was no hesitation, warning or small print telling me it might have invented the entire thing.

Then the system tried to open the record.

It was not there.

The agent had not found the wrong record. It had confidently quoted one that did not exist.

That should have taught me the lesson immediately.

It did not.

I learned that AI could be confidently wrong. What took longer was realising how often I was confidently wrong about the AI.

Confidence is part of the presentation

Language models are extraordinarily good at making an answer feel finished.

The sentences arrive cleanly. The structure is convincing. The answer does not look like somebody thinking aloud. It looks like somebody who knows.

That confidence belongs to the presentation, not the evidence.

I understood this in theory. In practice, I was still reading a polished answer and feeling that small internal click of certainty.

So I started asking other models.

Claude would tell me something. I would ask ChatGPT and Gemini. If all three gave me roughly the same answer, I felt more confident. If they disagreed, I knew I had found something worth understanding.

That was better than trusting the first answer.

It still was not proof.

Three models can agree because the answer is right. They can also agree because they were given the same bad information, made the same reasonable assumption or found the same plausible-looking rubbish.

Agreement is a signal.

It is not a receipt.

The green tick problem

I then made exactly the same mistake with the system itself.

At one point, a Slack bot returned a successful response. The technical step had worked. The system accepted what it had been asked to do.

I treated that as proof the notifications had been sent successfully.

The human-visible outcome was false.

Nothing useful had arrived where the person expected to see it.

The API had told the truth. I had asked it the wrong question.

It had proved that a service accepted a request. It had not proved that a human received, saw or could act on the result.

That distinction sounds painfully obvious once it is written down.

It was less obvious when the screen was green and I wanted the job to be finished.

A green tick proves only the thing that put the tick there.

Sometimes that is enough.

Often it is several miles away from the outcome you are claiming.

Done was another confident answer

The same pattern appeared again in the build.

Tests passed. Processes ran. Reports said complete. I had plenty of evidence that individual technical things were working.

Then live retesting exposed 10,583 silent failures underneath one of them.

The system had passed enough checks for me to call it done.

It had not passed the check that mattered.

Could it do the real job, in the live environment, in the way a real person would experience it?

I had confused evidence of activity with evidence of outcome.

Again.

AI makes this easier to do because it can produce an astonishing amount of supporting material. Plans. Test summaries. Audit reports. Completion notes. Neatly formatted explanations of why something should work.

You can be surrounded by evidence-shaped objects and still have no proof that the thing works.

The awkward part about verification

Generation is quick.

Verification is often slow and boring.

Opening the source. Finding the original record. Clicking the actual link. Logging into the live system. Looking at what the other person received. Trying the journey on the device it was meant for.

None of that feels as clever as building another agent.

It is also where most of the truth lives.

I began separating claims that I had previously bundled together.

The code exists.

The code runs.

The system accepted the action.

The action produced the intended result.

A human could see and use that result.

Those are five different claims.

Passing one does not quietly pass the other four.

This made me more annoying to work with, including for my own agents.

When one said something was complete, I started asking: complete where? Seen by whom? Compared with what? Can I open it? Can you show me the actual result rather than the report describing it?

The answers got less impressive.

They also became more useful.

I changed the question

I used to ask whether I trusted the AI.

That is too big a question to be useful.

I now ask whether I can verify this particular claim.

If it gives me a fact, where did the fact come from?

If it says it changed something, can I see the changed thing?

If it says a message was sent, did the person receive it?

If it says the job is complete, what would fail if that were not true?

Confidence still has a place. It helps me decide where to look next. It helps me move when certainty is impossible.

But it no longer gets promoted to proof because the answer sounds good or the box turned green.

The useful question became: what would still be true if this conversation disappeared?

The source would still exist.

The record would still open.

The person would still have the message.

The live thing would still work.

That change did more for the system than another model, another agent or another clever prompt.

Because by then, the thing I was really building was no longer the thing on the screen.