Somebody went and asked, on real people doing real work, and the finding is worse than you would guess.
In 2025, researchers at Microsoft Research and Carnegie Mellon surveyed 319 knowledge workers who used these tools at least weekly, and collected 936 first-hand examples of them doing it. They were looking for when people actually think critically about what comes back, and what makes them stop.
The headline result is one sentence:
“higher confidence in GenAI is associated with less critical thinking, while higher self-confidence is associated with more critical thinking.”
Read that twice. What they found is an association, in self-reported data, and I want to be exact about that before I build anything on it. They did not watch confidence cause the drop. They asked people, and the two moved together.
Here is what I think it describes, and you should treat this next part as my reading rather than their finding.
The better the tool gets, the more you trust it. The more you trust it, the less you check. The less you check, the less practised you are at checking, so the next time you are even less equipped to notice, and the tool by then is better again.
That is a trap with no floor in it. The survey is one turn of that loop, caught in a single moment, on 319 people. Whether it runs the way I have just described is the thing you can test on yourself, and this chapter is mostly about how.
Nothing in that loop contains a moment where something stops you. There's no alert. No amber light. The output does not look different on the day it is wrong.
That is the argument of this chapter, and it's the one I'd keep if you made me throw away every other technique in this book. Every other chapter is a method. This one is a habit, and habits are the things that decay when nobody is watching.
Here is where most writing on this becomes useless, because it only warns you about one direction.
The obvious failure is trusting it too much. You accept something that was wrong, it goes out with your name on it, and you find out later in a room you did not want to be in.
The failure nobody mentions is the opposite one. You reject something that was right. You spend an hour re-doing work that was already correct, you disregard an objection that would have saved you, or you decide the whole category is unreliable and go back to doing everything by hand while somebody two desks away produces four times as much.
Human factors research named both of these nearly thirty years ago, well before any of this. Parasuraman and Riley, in 1997, called them misuse and disuse: over-reliance on automation on one side, neglect of automation that would have helped on the other. The modern literature on working alongside these systems calls the same pair over-reliance and under-reliance, and it's consistent that both damage the quality of your decisions.
So this isn't a chapter telling you to be sceptical. Scepticism applied uniformly is just a slower way of being wrong, and it burns the hours the tools gave you back.
It's a balance, and the balance is the skill. Accept what's right. Test what you're unsure of. And spend the time you saved on the thing neither of you has thought of yet.
Left alone, this does not stay balanced. It drifts one way, and it's worth understanding why, because the reason is not laziness.
A well-formed answer removes the felt need to check it.
Look at what arrives when you ask one of these things a real question. It's structured. It's calm. It anticipates the obvious objection and deals with it in the third paragraph. It uses the vocabulary of your industry correctly. Every single one of those properties is a signal your brain has spent a lifetime reading as this has been thought about by somebody competent.
None of those properties has any connection to whether it's true.
That is the whole mechanism, and it's the same one as chapter three, arriving at a different point in the process. In chapter three it was a feeling of relief that moved your hand. Here it's the shape of the answer doing the same job. Both of them produce the identical outcome, which is that you move on.
Do that once and nothing happens. Do it for a year and something does.
Here is the discipline, and it is three moves in a fixed order. Not the three questions from chapter one · those interrogate an answer you have already got, and these decide what to do with it. It takes about ninety seconds and I would rather you did it badly than skipped it.
1. What here can I accept?
Start with acceptance, deliberately, because starting with doubt makes this exhausting and you'll stop inside a fortnight.
Most of what comes back is fine. The structure is fine. The gathering is fine. The arithmetic is usually fine. Accept it, out loud, and notice that you have made a decision rather than drifted into one.
The test for this is not does it feel right. It's would I have been able to produce this myself, and does it match what I already know to be true? Where the answer is yes to both, accept and move on. That is not laziness, it's the appropriate use of a thing that is genuinely better than you at that particular stage.
2. What am I unsure about, and what's the cheapest test?
Now find the parts you can't accept, and be specific. Not I don't fully trust this. Which sentence.
Then match the test to the stake, which is chapter five's whole argument arriving here in a different form:
- A figure you'd repeat in a meeting. Ask where it came from and go and look at the source. Two minutes. - A conclusion you'd act on. Make it argue the opposite, which is chapter seven. Twenty seconds. - Something you'd send to somebody who matters. Put it through a second system, which is chapter nine. Ninety seconds. - A claim you can't check at all. Say so, in the document, in your own words. That sentence costs you nothing and it's the one that protects you.
Notice that none of these is are you sure. Chapter seven explains why that question is worthless, and it's the question almost everybody asks.
3. What has nobody thought of?
This is the one that's actually hard, and it's the one worth your career.
The machine is genuinely good at half of it. Ask it directly what you have failed to consider and it will produce things you had not thought of, on any subject, reliably. Most people never ask, which is astonishing given how cheap the question is.
The other half is the one chapter two described · the absences nobody ever wrote down, which is precisely why nothing that reads can find them.
So run it in both directions, and the order matters. Ask the machine what is missing from the material. Then ask yourself what is missing from the situation. Doing it that way round means you are not competing with it, you are picking up where it stops, and you will be surprised how often its list makes yours obvious.
Ninety seconds is easy to say. Here is the whole of it on a single ordinary output, so you can see how little of it is work.
You asked for an analysis of why complaints rose last month. Back comes two pages: volumes by category, a rise concentrated in category three, three candidate causes, and a recommendation to review the escalation process.
What can I accept? The volumes, because they came out of the system and you can see the query. The arithmetic, because you spot-checked two rows and they hold. The observation that category three carries the rise, because you can see that yourself in ten seconds. That is most of the document, accepted deliberately, in about twenty seconds · and notice that you have now decided to accept it rather than drifted past it, which is the entire difference this chapter is about.
What am I unsure of, and what is the cheapest test? One sentence, and you can name it: the rise correlates with the new booking flow going live on the 8th. You are unsure because correlate is doing a lot of work in that sentence and you know the flow went live on the 8th because you were in the meeting.
Cheapest test, matched to the stake: this is going to your director and it will drive a decision about the booking flow, so it earns more than one pass. Show me the daily complaint counts for category three for six weeks either side of the 8th, and tell me what else changed in that window.
What comes back is that a second thing changed on the 11th. A supplier switch nobody had mentioned. And the daily counts move on the 11th, not the 8th.
What has nobody thought of? Run it both directions, which takes forty seconds. Ask the machine: what would explain this rise that is not in the material I gave you? It produces four things, one of which is seasonality you had not considered.
Then ask yourself, which is the half only you can do. And the answer is sitting there: complaints in category three are logged by the same two people who were on annual leave for the first week of the month. So the first week is not low. The first week is unrecorded, and the “rise” is partly a return to normal logging.
Ninety seconds, and the document changed twice. Once because the date was wrong, and once because the baseline was wrong · and the second one would have survived every check in this book except the one only you can run.
Richard Paul and Linda Elder spent a career on this and produced a framework used in universities everywhere. It has eight elements of reasoning and nine standards to judge them against, which is more than anybody remembers under pressure.
So take four of them, which is what actually survives contact with a Tuesday.
The purpose. What is this piece of work for? Not the task. The decision it feeds. Half of all bad output is a correct answer to the wrong purpose.
The question. Is the question it answered the question you needed answering? These systems are relentlessly obliging and will answer a nearby question beautifully rather than tell you yours was ill-formed.
The assumptions. What has been taken as given? Every answer rests on something unstated, and the unstated part is where the error lives. Ask for it explicitly: list the assumptions this depends on. It will.
The consequences. If this is wrong, who finds out, and when? That's chapter fourteen's question and it belongs here too, because the answer determines how much of this discipline the work deserves.
Four questions. Purpose, question, assumptions, consequences. You can hold that in your head in a lift.
Everything else in this book has a trigger. A brief gets written when work arrives. A score gets asked for when something looks finished. A second system gets opened when a decision matters.
This one has no trigger. Nothing in the design of any of these tools will ever say you have accepted eleven things in a row without checking one of them. The interface is built to be helpful, and a prompt to doubt it would be a strange feature for anybody to build.
Which leaves you. Not your judgement in the abstract, but a specific habit you decide to run.
Mine is the ninety seconds above, and I run it on anything I would put my name to. That is a low bar and it's deliberately low, because a discipline you actually keep beats a better one you abandon.
Yours might be different. What it cannot be is nothing, and I'll notice if something looks wrong is nothing, because the entire finding of that survey is that the better this gets, the less anything will look wrong.
---
### ▪ DO THIS > > Take the last thing you accepted without checking. Ten minutes. > > 1 · Find it. The most recent piece of output you used, forwarded or acted on without going back over it. There will be one from this week. > > 2 · Run the three moves in order. What can I accept. What am I unsure of and what's the cheapest test. What has nobody thought of. > > 3 · Then ask it directly: List the assumptions this answer depends on, and mark any you cannot support. > > 4 · Go and check exactly the ones it marked. Not all of them. The marked ones. > > Ten minutes the first time, ninety seconds after that. > > If you find nothing wrong: that's the good outcome and it's not a wasted ten minutes. You now have the right to put your name on it, which you did not have before you looked. > > You'll know it worked when you catch yourself doing move one deliberately, on something you would previously have accepted without noticing you had.
Ten chapters told you how to get more out of these tools. This one and the one before it are about the person doing it, because that is the part with no upgrade path and no support contract, and it's the only part your employer is actually buying.
Which covers a single exchange, thoroughly, from both ends. And almost nothing you are actually paid for is a single exchange.