← AIM · contents
Chapter 7

Make It Prove Itself

Are you sure?

Yes.

Of course it said yes. You asked the same system the same question in a slightly more anxious tone, and you gave it every signal that you wanted reassurance. It gave you reassurance. That isn't a check. That is a conversation with a mirror.

The last chapter was about noticing. This one is about what you do about it, and the honest news is that the obvious move doesn't work. Double-check that is worthless too, and so is is this accurate? All three produce the same output: a confident restatement of what it already told you, sometimes with a hedge bolted onto the front so it sounds humble.

What works is stranger and it takes about twenty seconds longer.

You have to give it a reason to disagree with you.

The mistake everyone makes first

Picture what you actually do; you have an answer you half-trust. You paste it back in and ask it to check its work.

Think about what you have just handed it. The conclusion, the reasoning that produced the conclusion, and your evident hope that the conclusion survives. It now has three pieces of information, and all three point the same way.

There is a word for asking someone to grade work while showing them the answer they already gave and looking hopeful. It isn't a check. It is a formality.

The whole of this chapter is variations on one move: change the thing's incentives before you ask it anything.

Make it solve the problem before it reads yours

Start here, because it's the simplest version and it works immediately.

We built a study tool for a teenager sitting exams. One system produces an answer to a question. A second system marks it. That much is obvious.

The instruction to the marker is the part that matters, and it's one sentence:

First solve the question yourself, from scratch, before you read the proposed answer properly. Do not anchor on it.

That is it. That is the whole technique.

Without it, a marker reads the answer, finds it plausible, and confirms it. Plausible is a very low bar and almost everything wrong is plausible, otherwise nobody would have written it down. With it, the marker arrives at the problem with its own answer already formed, and now there are two answers to compare rather than one answer to approve.

You can do this on anything. A forecast, a summary, a recommendation, a price. Work this out yourself first. Then read what I have and tell me where we differ.

The difference in output is not subtle. You stop getting this looks correct and start getting I get a different number, here is where it diverges.

Here is the whole thing on an ordinary piece of work, because it is thirty seconds of typing and almost nobody has seen the two sitting next to each other.

Picture a cost estimate for moving a team of twelve into a new office · invented for this example, so treat every figure in it as illustration rather than as one of mine. Twelve desks, the fit-out, cabling, a month of overlap on both leases, and the removal itself. The spreadsheet says £84,000 and it is going to the finance director on Thursday.

You already know what happens if you paste it in and ask whether it looks right. The structure is logical, the major categories are covered, the contingency seems appropriate, and you might want to confirm the cabling quote. Courteous, accurate, useless · because it read your arithmetic before it did any of its own, and from that moment there was only ever one answer available to it.

So do not send it. Send this instead:

I am moving twelve people into a new office. Fit-out, desks, cabling, one month of overlap on both leases, and the removal. Before you read anything of mine, build the estimate yourself from scratch and show your assumptions. Then I will send you mine and you tell me where we differ and which of us is wrong.

It comes back at £108,000, and the interesting part is not the gap. It is the four lines underneath it, because two of them are things you had thought about and two of them are not.

It has assumed dilapidations on the old lease. You had not, because your landlord has never mentioned it and you have not read the clause. It has assumed a week of reduced productivity across twelve people and priced it, which you would probably argue with, and you should · but you will argue with it out loud, in front of the finance director, which is a much better position than being asked about it and having nothing.

And then the one that pays for the whole exercise. It has your overlap at one month and flagged that a fit-out of this size slipping by two weeks is the ordinary case rather than the bad case, so the overlap is the line most likely to move and the one you have least control over.

That is not a checking result. That is somebody thinking about your problem before they thought about your answer, and you got it by withholding your answer for ninety seconds.

Now send yours and ask where you differ. What comes back is a list you can act on, and the £84,000 either survives with reasons attached or it does not, and both of those are better Thursdays than the one where a courteous system told you the structure was logical.

And when they agree, that agreement means something, because it was arrived at twice from two directions rather than once and then nodded at.

Give it a different job, not a different tone

The second move is to stop asking it to check and start telling it who to be.

There is a piece of our system whose entire job is to attack a proposal before it goes anywhere; it is cast as a diligence analyst at a major bank looking at an acquisition. The instruction includes, in plain words:

Your job is NOT to be a cheerleader.

And it isn't advisory. If it returns anything marked critical, or more than two marked high, the thing doesn't proceed. The gate is code, not judgement.

Notice what has changed. Nobody asked it to be more careful. Careful is a temperature setting and it does nothing. It was given a role with different incentives, and roles carry behaviour that instructions don't.

A diligence analyst who approves everything is bad at their job. That is the load-bearing part. The role itself supplies the motivation to find something, so you no longer have to keep asking.

You can do this in one sentence with no system at all. You are the person in the room whose job is to stop this going out. What stops it?

Or, from the last chapter, four words that carry the whole posture: use an adversarial agent. It knows what that means.

There is a smaller version of the same idea running elsewhere in our estate, checking outbound writing for anything that could be read as defamatory. It is cast as a cautious media lawyer, and the reason it is cast that way is written into the code as a comment:

A false positive costs 3 seconds. A false negative costs a relationship.

That is the calculation to make before you decide whether a check is worth running. Not how likely is this to be wrong. What does each kind of wrong cost me? When those two numbers are that far apart, you do not need the check to be right very often for it to pay.

Starve the checker

Now the least obvious idea in this chapter, and the one I would most like you to take.

Give the checker less information than you gave the worker.

Every instinct says the opposite. More context, better judgement — that is true when somebody is doing the work. It is false when someone is checking it, because most of that context is the reasoning you are trying to test, and handing it over is handing over the answer.

We do this deliberately in four places and each one deprives the checker of something specific.

The exam marker is not allowed to read the proposed answer before forming its own.

A quality reviewer that assesses whether a piece of client correspondence did its job is given only the original inbound message. Not the brief, not the strategy, not what we were trying to achieve. It has to work out what the job was from the client's own words. If it can't, we didn't understand the client.

A confidence challenger, whose job is to push back on how certain a claim is, is given the claim and a count of how many sources support it, and never the reasoning that produced the confidence. Its instruction says so directly:

You are NOT the analyst who made this claim. Your incentives are opposite: they want high confidence, you want accurate confidence.

And a set of assessment lenses is deliberately run blind, without knowing whose interests they're supposed to be serving, so that the assessment isn't quietly bent towards the answer somebody wanted.

The pattern is the same every time. Whatever would let the checker shortcut to your conclusion is exactly the thing you withhold.

And the counter-example, which is why this is a chapter

Now the part that stops this being a neat rule you apply everywhere and get burned by.

We took the same idea and applied it in the wrong place.

We had a grader, whose job was to score drafts on how well they read the situation. And in the spirit of blindness, it was run without the original message the draft was responding to.

The results are recorded. In seven of twenty audited items, the grader produced some version of the same sentence:

Cannot assess decode accuracy without seeing the original inbound.

It couldn't do the job. It said so. And because it still had to return a number, it returned low ones, capping the average at 31.9 out of 100.

Every one of those scores was garbage, and they looked exactly like real scores. A number arrived, it was in range, it went into the average. Nothing anywhere flagged that a third of them were the machine telling us it had been given an impossible task.

Here is the distinction, and it's worth memorising because it is the difference between the technique working and the technique poisoning your data.

Starve the checker. Never starve the grader.

A checker asks is this right, and it needs to be prevented from seeing your answer.

A grader asks how good is this, and it needs the standard to measure against. Take the standard away and it doesn't refuse. It guesses, and hands you a number that looks identical to a real one.

Blindness is a tool. It isn't a virtue. Ask yourself which of the two you're building, every time, and if the answer is both then you need two of them.

When you have only got one of them

Everything above assumes you can reach a second system, and the redaction story in the last chapter is the argument for why you should. I am not going to run those numbers past you twice.

Here is the question that chapter left open. It is Tuesday, you have one tool, your employer has approved exactly that one, and a second opinion from a different company is not available to you today.

Then change the role instead of the model.

It is the weaker version and I want to be straight about how much weaker. A different model brings a different set of assumptions, and assumptions are the thing you are actually testing. A different role brings the same assumptions with a different attitude on top of them. It will not catch what the system simply does not know.

What it does catch is the far commoner failure, which is a system agreeing with you because agreeing was the path of least resistance. A diligence analyst who approves everything is bad at their job. A helpful assistant who approves everything is doing exactly what it was asked. Same engine, opposite incentive, and the incentive is most of what you were buying.

So use the role when the model is all you have, and know what you have bought: protection against easy agreement, not against a shared blind spot. The second is why chapter nine exists.

The three things to type

Everything above collapses into three instructions. None needs any tooling.

1. Work this out yourself before you read my version. Then tell me where we differ.

2. You are the person whose job is to stop this. What stops it? List every reason this fails, worst first.

3. Argue the opposite of your own conclusion, as convincingly as you can. Then tell me which argument is stronger and why.

That third one is worth a note. We run a version of it on a system that produces group assessments, and it measures how much the voices agree with each other. When agreement gets too tight, it does something deliberate: it takes the most analytical of the voices and instructs it to

argue the OPPOSITE case regardless of your personal assessment.

Because unanimity isn't evidence. Unanimity is very often the sound of everyone reading the same thing and finding it plausible. A room that always agrees isn't a room that has checked anything, and this is as true of people as it is of software.

Confidence is not evidence, and the two are not even related

Now the most important idea in this chapter, and it is the one that will save you from the worst mistake available to you.

These systems produce confidence and accuracy almost independently. Not loosely coupled. Almost independently, and the gap is far too wide to lean on.

The clearest demonstration I know is a published result that our own confidence gate cites in its code, as the reason it exists at all. On a reasoning benchmark, a system reported an average of 89 per cent confidence across five problems.

It got zero of the five right.

Read those two numbers together until they sit properly. Not one wrong out of five. Not three. All five, at eighty-nine per cent confidence, with no internal signal whatsoever that anything had gone wrong.

There is no mechanism in there that connects how sure it sounds to how right it's. It has never had one. Confidence is a writing style.

So we stopped asking it how confident it was and started counting instead.

The rule is arithmetic and it overrules the model's own judgement no matter how convincing the argument:

Distinct sourcesMaximum confidence allowed
None35%
One or two50%
Three or four75%
Five or more95%

If it wants to claim ninety per cent on something with no sources behind it, it does not get to. The ceiling is thirty-five and the ceiling is code.

You do not need our system to use this. Count the sources yourself. It takes ten seconds. Zero sources isn't a weak claim, it is a different kind of object, and treating it as a fact you can act on is how people end up standing in front of a room defending something that was never checked.

There is one more piece of that gate worth stealing. A second system is brought in specifically to argue the confidence down, and when its verdict differs from the original, the original is kept alongside it rather than overwritten. You can see both. If the challenger inflates confidence by more than twenty points, the whole thing is rejected outright, because a challenger that talks you up isn't a challenger.

Keep the first answer. The disagreement is the useful part.

So I went and looked

I have described a set of mechanisms for making things prove themselves. They are real, they're running, and I have shown you the code comments.

Then I did to them what this chapter has spent twenty pages telling you to do to everything else. I stopped asking whether they looked right and went and counted.

Here is what came back, and it is worse than not knowing.

Start with the simple part. The place where a challenge result is supposed to accumulate, so that I could look back and say the challenger overturned the original answer in X per cent of cases, holds zero records. Not thin data. None. I checked every copy of it I could find across our systems, in case the results had landed somewhere I had forgotten about, and they had not.

My first assumption was that somebody had forgotten to switch on the logging. That would have been the comfortable answer and it is not the one.

The challenger has never run. One of the two things that triggers it is a flag, and nothing in the system has ever set that flag, so the queue it reads has been empty since the day it was written. It has been sitting there the whole time, correct, complete, and waiting for work that could not arrive.

Then I looked at the live path, which does not depend on that flag, and this is the part I would like you to hold on to.

It fires when two conditions are met at once. Agreement has to be tight, and the spread of opinion has to be narrow. Two safeguards, and I had read that line more than once and thought it looked careful.

They are the same condition. Both of them measure how far apart the votes are, one as a spread and one as a statistic derived from the spread, and the first one is written so loosely that it was satisfied on fifty-eight occasions out of fifty-eight. It has never once been false. So a gate I believed was two independent tests is one test wearing two coats, and it took me four minutes with a calculator to find that out, having never once thought to look.

And one more, because it is the detail that stops this being abstract. The condition that does the actual work triggers when the spread is fifteen or below. Fifty-two of those fifty-eight sit at seventeen or eighteen. The entire working population of this thing is clustered two or three points above the line at which it does anything at all.

So the honest position on this chapter is not that I built an adversary and forgot to measure it. It is that I built an adversary, wired it in, admired it, and it could not have fired. Every case I have shown you above is real and every one of them was produced by a different mechanism than the one I was proudest of.

The repair is specified and handed to the people who own that code, with the arithmetic attached so nobody has to find it twice.

There is something almost funny about that, and it is worth sitting with rather than laughing at. The easiest system to forget to check is the one whose whole job is checking. It looks like diligence. It reports success. Nobody audits the auditor, because the auditor is the thing that was supposed to stop this happening.

What this is actually for

Step back from the mechanics, because the mechanics aren't the point and I don't want you to leave this chapter with three prompts and nothing else.

What you're building isn't accuracy. Accuracy is a by-product.

You are building the ability to say, out loud, in a room where somebody is pushing back: I know this is right, and here is how I know. Not the system said so. Not I checked it. Here is the thing I did that would have caught it if it were wrong, and it did not catch anything.

That sentence is rare. Most people can't say it about their own work, and they can't say it because they've never built anything that could have caught them.

And you have just watched what it costs to be able to say it honestly. I can tell you the tenement figure is right, and I can name the check that found it. I cannot tell you the adversarial gate has ever caught anything, because I went and counted and it has not. Both of those sentences come out of the same habit, and you do not get to say the first one unless you are willing to say the second.

That is the whole trade. Somebody who only tells you what worked is telling you about half of a process, and you have no way of knowing which half.

It is also the sentence that separates the two people I keep coming back to. One of them uses these tools and produces work that's usually fine. The other one can stand behind it. Only one of those is being paid for judgement, and judgement is the only thing in this whole field that's not getting cheaper.

So make it argue with you. Give it a reason to. Withhold what would let it agree with you cheaply, and never withhold what it needs to be fair.

And then, when it has fought you and lost, you'll have something you didn't have before, which is a number you would put your name on.

Which raises the question I've been avoiding for two chapters, because a number you would put your name on isn't the same as a number that's good enough, and nobody has ever told you where that line is.

Watch this chapter
A short illustrated film of Chapter 7. The narration is the author’s own words from the chapter. The animation, the voice and the score are AI-generated.
3,665 words