Here is a sentence that went out to a client, in a document with my name on it.
133 tenements sit inside this frame, and three companies hold all of them.
Read it again. It is a good sentence. It is confident, it is specific, and it does the thing a client is paying for, which is to turn a mess into a picture they can act on. Three companies. You know who to call.
The real number was 418 tenements, held by 79 different companies. The largest of them held under 10 per cent. Two of the three companies I named held nothing at all inside that boundary.
Not slightly off. Wrong by a factor of three on the count, and wrong in a way that inverted the advice. Walk up to one of three companies is a strategy. Walk up to one of 79 is a different job entirely, and the client would have made a plan on the first one.
I want to be careful about what this chapter is, because the obvious reading is the wrong one.
This isn't a chapter about catching mistakes.
Most of the time when I go back and ask again, the first answer wasn't wrong.
It was correct, and it was not right. It was looking at the thing from an angle that did not help. Or it was explained in a way that would have lost the person who has to read it. Or it was three paragraphs of prose when it needed to be a table, and no amount of rewriting those paragraphs was ever going to fix that.
That is the ordinary case. Accuracy is the exception.
I am starting with the tenement story because it's the sharpest example and because the stakes are visible. But if you take away from this chapter that the job is fact-checking, you'll do the fact-checking, feel diligent, and still hand people work that does not land.
The job is bigger and it's more interesting. The job is noticing that you have been handed the wrong shape.
You already have the signal; nobody has told you it is a signal.
It sounds like this.
I hope this is right.
Or: I didn't know that. Or: that isn't what I thought it would be. Or something wordless and half a second long that you would struggle to write down at all.
You have felt it. Over a forecast, a summary, a quote you were about to send a customer. It arrives, you notice it, and you move on, because the answer looks fine and there are nine more things on the list.
Stop moving on. That is the whole technique, and everything else in this chapter is what to do in the thirty seconds after you stop.
Let me be honest about the hit rate, because a book that tells you the feeling is always right is selling you something.
Sometimes I get the feeling, I go and check, and the original answer holds up perfectly. That happens. It isn't a wasted six minutes, and I'll come back to why at the end of this chapter, because the reason is the most important thing in it.
But generally, when that thought turns up, something is off. And generally what is off is not the fact. It is the framing.
I want to give you the actual reason, because it's a rule you can use tomorrow and it will feel wrong at first.
133 is specific. Specific means I want proof, and proof should be easy.
Think about what your instinct does with numbers. Someone says roughly a third of the market and you file it as an estimate. Someone says 31.4 per cent of the market and you relax. The decimal point reads as diligence. Somebody counted.
That instinct is backwards, and it's expensive.
A round number is honest about being approximate. A specific number is making a claim about where it came from. It is saying: I did not estimate this, I counted it, and the count exists somewhere. If that somewhere cannot be produced in about ten seconds, the specificity was decoration. It was the shape of rigour without the substance.
So the more precise the number, the faster you should be able to see its source, and the more suspicious you should be when you cannot.
One hundred and thirty-three isn't a number anyone guesses. It came from somewhere, and I went to look at the somewhere.
The somewhere turned out to be one zoom level of a commercial mapping tool.
The count was real. It was a statewide total for a handful of holders, sitting on a viewport at a particular zoom. Somebody read it off the screen and used it in prose two paragraphs later as a local concentration claim. The statistics card at the top of the page carried the caveat about what the number covered. The prose below it dropped the caveat and kept the number.
Nobody lied. There was no moment where anyone chose to mislead. A true statement about one thing became a false statement about another thing, on the same page, two paragraphs apart, because a caveat didn't travel.
That is the ordinary way a true claim becomes a false one, and no amount of good intent prevents it. Only a check does.
But there was a second layer, and it's the one that should frighten you, because it's the one you cannot see.
The tool we used to query the government tenement register was broken in four ways. Three of them were the sort that announce themselves. The fourth didn't.
The register's data service returns its field names in lower case. Our code was reading them in upper case. Every genuine result came back, was processed correctly, and normalised into rows of perfect nulls.
Sit with what that means. A real result looked exactly like no result. Not an error. Not a warning. Not an empty response you might question. Rows of clean, well-formed, entirely empty data, produced by a system reporting complete success.
We had seen this before and not understood it. There had been an earlier oddity where a query over a known operating mine returned a count of zero, and it had been filed as strange rather than as diagnostic; it was the same bug, showing us its face, and we didn't recognise it.
When it was repaired and pointed at the same ground, it returned 418.
Here is what I didn't expect, and it is the reason I would run this check even when the first answer is fine.
The correct number was not just less wrong. It was worth more.
Four hundred and eighteen tenements across 79 holders, with no single company dominating, sounds like worse news for a client who wanted a simple picture. But the same query that produced it also produced this: 73 of those tenements expire within twelve months. 127 within 24. 44 more are already past their expiry date and still sitting on the register.
Expiry is the opening. Ground that is about to come free is the single most actionable thing you can hand somebody in that industry, and it is invisible from a map. It only exists in the register.
The wrong number wasn't merely wrong. It was hiding the answer.
I have found this again and again and it's the argument of this chapter. The check is sold to you as insurance, a cost you pay to avoid a loss. That is not what it is. Most of the value in the second look is not the error you avoid. It is the thing you find while you're looking.
I am going to give you the arithmetic, because this is where most people decide the whole idea is too expensive and they're wrong by an order of magnitude.
One minute to think about what I actually doubted. Not the whole document, the specific claim.
Two to five minutes to write the instruction properly, so the check would go to the source rather than re-asking the same system the same question.
Fifteen minutes while it worked. I wasn't in the room. I was doing something else.
Then again, with a different instruction, and a different model.
Six minutes of my attention. That is the real cost, against a figure that was wrong by a factor of three in a document going to a client with my name on it.
The fifteen minutes of machine time doesn't count and I want to say plainly why, because getting this backwards is what stops people. The scarce resource isn't compute. Compute is close to free and it is getting cheaper while you read this. The scarce resource is your judgement, and the check spends almost none of it.
Six minutes. If you take one number from this book, take that one.
The instinct, once you accept all this, is to ask again. Same question, more forcefully. Are you sure?
That is the least useful version of this and it produces a reliably worthless answer, because you have asked the same mind the same question and given it a reason to agree with you.
Three checks, and each one is a different job.
The first gets the answer. That is the part you already do.
The second attacks it.
You don't need clever wording. You say the words use an adversarial agent and it knows the posture you want. That phrase carries the whole instruction.
Then you make it show its work, and be specific about the form:
Give me the sources as hyperlinks I can click. Put the claims in a grid, one row per claim, with the source next to each. Add a column rating your certainty on each row from 0 to 100.
The 0 to 100 column is the part people skip and it is the part that changes what you can do next. An answer without it arrives flat, every sentence carrying the same apparent weight, and you have to read all of it with equal suspicion. An answer with it arrives ranked. You look at the rows scoring 40 and you know exactly where the six minutes goes.
Scores are how you work with any of this. When it compares options, ask it to grade them. When it makes a recommendation, ask what the alternatives scored and why. The number isn't the point. The number is what lets you push on the reasoning, and an opinion you can't push on is worth very little.
There is one more thing to put in that second instruction and almost nobody does it.
Tell it who is going to read this.
Not the topic. The person. Their role, what they already know, what they will do with it, what would waste their time. It changes the sophistication, the vocabulary, the length, and what gets left out. It is the highest-leverage sentence you can add to any instruction and it costs nine words.
And a diagnostic that comes free with it: if you're not sure whether it understood, ask it who it assumed the audience was. If the answer surprises you, you've just found out why the output felt wrong. That is often the entire mystery solved in one question.
The third makes it land.
This is the one nobody reaches and it's where the career value is.
It isn't enough to have the right information. The person receiving it has to understand it, and they have to be able to run their own decision-making process on it. If they cannot, you haven't finished. You have just moved the work to them and attached your name to it.
So the third pass is not another verification. It is: show me this differently. Make it visual. Make it a table. Explain it for somebody who has ninety seconds and a decision to make.
People are not the same. Some will take an audio summary before a meeting and have everything they need. Some want to read a page and be done. Some have to sit with it themselves and turn it over. You aren't going to know which one you've got, so the third pass is where you produce the version that survives all three.
And there is a specific moment you are aiming for, which chapter sixteen is entirely about. It is not the moment you sent it. It is the moment they could see it, and those are almost never the same moment.
Most people stop at the first check. A few reach the second. The third is where you stop being someone who produces correct work and start being someone whose work gets acted on, and those are different reputations that get paid differently.
I have been describing this without showing it, which is the wrong way round for a chapter about what to type.
So here is the exchange, reconstructed from our own commit record and the correction notes rather than from a screen capture — the instructions are the ones we use, and the numbers are the ones that came back.
The first ask, which is where almost everybody stops:
Summarise the mining tenure across this area for a client report. Who holds the ground?
Back came the paragraph that went into the document. 133 tenements, three companies holding all of them, a clean picture and a clear recommendation.
The second ask. Note that it isn't are you sure, and note that it names a source rather than asking for reassurance:
Use an adversarial agent. Do not re-use anything from the previous answer. Query the state tenement register directly, across every permit layer, for this boundary. Give me the result as a grid, one row per holder, with a link to the register entry and a certainty score out of 100 on each row. Tell me explicitly what you could not verify.
That returned 418 tenements across 79 holders, and; this is the part that mattered — two of the three companies named in the first answer held nothing at all inside the boundary.
The third ask, which is the one nobody reaches:
The client is a landholder, not a mining analyst, reading this before a meeting. Show me the same finding in the form that makes the decision obvious to them. What is the one number on this page?
That produced the expiry table. 73 tenements coming free within twelve months, which is the commercially useful fact and doesn't appear anywhere in the first two answers.
Look at what changed across the three. The first asked for a summary. The second changed the source. The third changed the reader. Only one of them was about accuracy, and it's the one that gets all the attention.
⚠ One honest note. I am reconstructing these from the record rather than pasting a session log, because I didn't keep one. That is a small failure of exactly the kind this book keeps admitting — the most instructive exchange of the year, and no transcript of it. Keep yours. They are worth more than you think, and you will want one the day somebody asks how you knew.
I want to give you the sharpest version of why the check has to come from somewhere else, and for that I need a different story.
We built a system whose job was to find sensitive information in text and hide it before it went anywhere. Names, account numbers, medical details, that sort of thing. Getting it wrong in either direction is bad. Miss something and private data leaves. Hide too much and the output is useless.
So it was tested properly. Sixty-nine hand-built cases across fifteen categories, positive and negative, with a harness that pulled its patterns directly out of the shipped code so the test could never drift from the running system.
It scored a perfect one point zero on precision and a perfect one point zero on recall. Everything it should have caught, it caught. Nothing it should have left alone, it touched.
Then somebody wrote a test set specifically designed to defeat it.
Precision fell to 0.888. 31 false alarms. Thirty things it missed entirely.
Nothing about the system had changed. The only thing that changed was who wrote the test.
That is the whole lesson and it applies to every piece of work you'll ever check. Your test set was built by the same mind that built the thing it tests. It tests what that mind already thought of. It cannot test what that mind did not think of, and what that mind did not think of is precisely the thing that's going to hurt you.
The number that means anything is the one produced by something trying to break you.
One targeted fix afterwards took the false alarms from 31 down to nine and precision to 0.966. So the adversarial test did not just measure the problem; it was the only thing that could have found it.
This is also, exactly, why the second and third checks need a different instruction and a different model. Not the same system asked twice in a nicer tone. A second opinion from the same mind is not a second opinion. It is the same opinion, with more confidence attached.
One more, shorter, and it's the one I think about most.
A document we published told people that the encryption protecting their data used 600,000 iterations of a particular key-strengthening function. It is a security parameter. Higher is stronger. 600,000 is a good number.
I opened the file. The shipped code reads 100,000.
A six-fold gap between what the document said and what the software did. No one lied. At some point the number in the code changed and the number in the document did not, and everybody involved continued to believe a thing that had quietly stopped being true.
It was found by reading one line of each.
I am telling you because a chapter about checking things, written by somebody pretending he had checked everything, would be worth nothing to you.
Here is why it matters more than it looks.
To whoever finds that gap, it doesn't read as documentation rot. It reads as a lie. They do not have access to the meeting where nobody decided anything. They have a document that says one thing and a system that does another, with your name on the document.
And this is where the stakes stop being about accuracy at all.
You would be far more forgiving of a person than you'll ever be of a machine.
A colleague gets a number wrong in a deck and you think: bad week. Same number, same deck, and it emerges the machine produced it and nobody checked, and it doesn't read as a bad week. It reads as carelessness with a system you did not respect enough to supervise.
I don't know that this is fair. I know that it is true, and true is what you have to plan around.
So understand what you're actually protecting, because it's not the accuracy of a document.
It is treated as something you claimed. As a lie. And other people find out.
The loss of trust from one wrong figure is enormous and it is not proportionate to the size of the error. It attaches to everything else you've given them, including the parts that were right. People don't go back and re-audit your previous work when they find one bad number. They just quietly discount all of it.
Which means the thing you're building, when you build the habit of checking, is not a reputation for accuracy. It is a reputation for being someone whose work does not need checking. That is a much rarer thing and it's worth a great deal more.
And this is the difference. Not between someone who uses these tools and someone who does not, because soon that will be everybody. The difference is between a professional who can be trusted with this and someone who just uses it. One of those gets paid.
I owe you the other side, because everything above pushes in one direction and a technique with no cost is a technique somebody is overselling.
We had a line in our outreach that worked. It closed messages warmly, it tested better than anything else, and the internal write-up celebrated it as the thing that made people reply.
The line was either way, rooting for you.
In Australian English, rooting is crude slang. It isn't a subtle regionalism. It is the kind of word that ends a conversation with somebody you were trying to build a relationship with.
It went out. In a real message, to a real person, in April.
It had been caught. Four separate checks in our system knew about that word. The one check sitting at the point of sending did not, and that is the only one that had to fail.
So: the version that tests best can be the version that ends the relationship. Iterating towards what performs will find you things that perform, and performance isn't the only thing you're optimising for. Somewhere in the loop, something has to be asking a question that has nothing to do with whether it works.
There is a second cost and it is quieter. If you check everything, you check nothing, because you won't sustain it and you'll start waving things through to catch up. The signal is what makes this affordable. You are not checking all your work. You are checking the parts where the sentence turned up in your head, and that's a small number of things per week.
Which brings me back to the case I left open.
Sometimes I get it, run the checks, and the first answer was right all along.
Six minutes, no error found, nothing changed.
That is not a failed check and I want to be precise about why, because it is the most useful idea in this chapter.
What you bought was not a correction. What you bought was the right to put your name on it. Before the check you were about to send something you privately hoped was right. After it, you're sending something you know is right. Those are different objects, and only one of them lets you stand behind it when somebody pushes back in a meeting.
There is no version of this where you get that feeling and are better off ignoring it. Either you find something, and you've saved yourself. Or you find nothing, and you've converted a hope into a certainty for the price of six minutes.
The only losing move is the one everybody makes, which is to feel it and carry on.
So start there. Not with a system, not with a process, not with anything you have to remember. Just stop overriding the sentence in your own head.
Everything after this chapter is about what to do once you've stopped, and the first of those things is the one I've been circling since the tenement number.
Because noticing that an answer might be wrong isn't the same as being able to make it prove that it's right, and the second one is a thing you have to build.