← AIM · contents
Chapter 9

Use More Than One

Open a second tool. Put the same question in. Look only at where the two answers part company, and spend your attention there.

Ninety seconds, and that's the entire method. The rest of this chapter is why it works, and why the people who most need it are the ones least likely to do it.

Ask yourself a question about a hospital.

Does a patient on a ward need a surgeon to check on them? Yes. Obviously yes. There are decisions on that ward that only a surgeon can make, and getting them wrong isn't recoverable.

So does the surgeon do every check?

Of course not. Nurses check on that patient far more often than the surgeon ever will. They take the observations, they watch the trend, they know what normal looks like for this person today. And when something moves, they call the surgeon back in.

Nobody thinks this is a compromise. Nobody thinks the hospital is cutting corners. It's simply how you run a ward. A surgeon doing hourly observations on forty patients would be ruinously expensive and, by about hour six, worse at it than the nurses.

That is how you run these tools, and almost nobody does it.

The person sitting next to you has one. They use it for everything. The hard analysis and the tidy-up of a paragraph, the strategic call and the reformatting of a list. They are paying surgeon rates to take somebody's blood pressure, and they're getting worse answers for more money, because the expensive one isn't better at everything. Judgement is where it earns its money, and most work isn't judgement.

What we actually do

Twenty-five different jobs run across our systems. Each one has a written policy about which tool handles it.

Eighteen run on a cheap, fast model. Six are pinned to the expensive one. Which leaves one I cannot place from my own table, and I have left it there rather than tidy the count.

The ratio isn't the interesting part. The interesting part is that every single pinning carries a written reason, in plain language, sitting in the file next to the setting. Two of them read:

IP-sensitive outreach, quality non-negotiable.

Acquisition decisions, Opus tier required.

There's the discipline, and you can adopt it this afternoon without any of our software.

If you cannot write down why this particular job needs the expensive one, it does not.

Try it on your own week. Most of what you do won't survive that sentence, and the things that do survive it are the things you should be spending your attention on anyway. The exercise is not really about cost. It is a way of finding out which of your work is judgement and which of it's typing, and most people have never separated the two.

The nurses report back

The part of the analogy that does the real work is the last part, and it's easy to skip.

The nurses don't just do the cheap tasks. They report back. They gather the observations, they notice when something is outside the normal range, and they escalate. The surgeon's expensive attention is spent on the escalation, not on the gathering.

Build that shape. Cheap and fast for collecting, listing, extracting, formatting, first passes, checking things against a rule. Expensive for the judgement call that arrives at the end of all that, holding everything the cheap ones brought back.

A researcher gathering forty documents doesn't need to be brilliant. A person deciding what those forty documents mean does.

Set it up the other way round, which is what happens by default when you use one tool for everything, and you spend your best resource on the gathering and arrive at the decision with nothing left.

Two of them disagreeing is the point

Now the other half of using more than one, and it has nothing to do with cost.

One system agreeing with itself is a single point of failure with a confident voice.

You saw this in chapter 6 from one direction: a checker that has seen your answer will confirm it. Here is the other direction. Two genuinely different systems, given the same problem, produce genuinely different answers, and the gap between them is information you cannot get any other way.

When they agree, you have something. Not proof; they aren't independent, they read much of the same internet and fail in correlated ways. But two routes that don't share your framing beat one route travelled twice.

When they disagree, you've found the thing worth your attention, and you have found it in about ninety seconds. That disagreement isn't a problem to resolve before you can get on with the work. It is the work. It is the map pointing at the one paragraph where the difficulty actually lives.

Which is what the previous two chapters were driving at when they kept saying a different model, not just a different question. Rephrasing gets you the same assumptions back. Only a different system gets you different ones, and assumptions are the thing under test.

Nothing about this requires infrastructure. Open a second tool. Paste the same question. Look at where the two answers part company, and spend your six minutes there.

Which two

I have written a chapter called Use More Than One and not yet told you which ones I use, and you would be entitled to read that as evasion, because in most books it is.

So: as I write this, in the middle of 2026, I use ChatGPT and Claude.

Now the caveat, and it is the reason most books on this subject will not name anything at all. By the time you are holding this, at least one of those names will have moved. A version number will have changed, a capability will have arrived that reorders which one I reach for first, or one of them will have been renamed by a marketing department. That is not an argument for saying nothing. It is an argument for dating the claim, so that you can see how old it is and discount it yourself, which is exactly what chapter sixteen asks you to do to your own work.

My pair matters less than whether yours are genuinely two, and this is the practical trap. A great many products that look like competitors are running the same engine underneath, licensed from the same handful of companies and dressed by different marketing departments. Your bank's assistant, your CRM's assistant and the tool your team bought last year may all be one system with three logos, and comparing them will feel exactly like a second opinion while giving you none.

So find out what is actually underneath before you trust any comparison. It takes one search.

Then stop taking my word for the pairing. Put the same real question through whichever two you can reach, on a decision you have to make this week, and watch which pair disagrees in ways that help you. That costs an afternoon, it is the only test that answers this for your job rather than for mine, and you finish it having picked your tools on evidence instead of on somebody else's habit.

One caution, because a second tool is a second place your work now lives. Whatever your employer approved, it probably approved one system and not two. Chapter fourteen shows how to strip the part that matters before it leaves, which makes this a question of habit rather than permission.

The reasons matter more than the ratio

I gave you eighteen and six. Do not copy the ratio.

Our eighteen and six comes from the shape of our work, and yours will be different. What transfers is that every one of the twenty-five has a reason attached that a human wrote.

None of that's bureaucracy. What it buys you is a decision you can review later. By you, once you've forgotten why you chose it, or by whoever inherits it. A setting with no reason attached is indistinguishable from an accident, and six months later nobody can tell whether it was a considered choice or a default nobody looked at.

Apply the same rule to your own habits. Always reaching for the expensive one on a certain kind of task? Write down the sentence explaining why. Can't? You've found something to change. Can? You now have a decision you can defend when somebody asks why the bill looks like that.

Know when you paid without deciding to

A quieter failure, and one you won't notice for months.

Systems fall back. Something is unavailable, something times out, something is rate-limited, and quietly the work routes to the expensive path instead. The output is fine. Nothing is broken. Nothing tells you.

We found ours the ordinary way, which is that the numbers looked wrong. The note in the code afterwards is blunt:

Silent fallback was costing real dollars on calls designed to be free.

Every fallback is now counted and classified, so a route designed to be cheap that has been quietly running expensive for a fortnight shows up as a number rather than as a surprise at the end of the month.

Your version of this is smaller and it's the same shape. Know which tool actually answered. Not which one you meant to ask. If you have a paid tier and a free one, or a fast setting and a slow one, check occasionally that what you are getting is what you chose. The gap between intended and actual is where money and quality both leak, and neither leak announces itself.

The counter-example, because I have made this sound too clean

I have spent this chapter telling you to be disciplined about cost. Here is what happened when we were.

We put in a cost-control gate. Sensible thing. It capped daily spend, it capped individual tools, and one of those caps was set to five dollars.

The gate did its job. It blocked automated calls once the limit was reached.

And a feature people were actually using stopped working. Users saw a failure. They didn't see this is temporarily unavailable to control costs, they saw the thing break.

Saving money broke the product.

The lesson isn't that cost control is wrong. The lesson is that a limit with no visible failure mode is a trap, and the person who set it's never the person who discovers it.

So when you build any rule for yourself about which tool to use, ask what happens at the boundary. When the cheap one isn't good enough for this particular case, how do you find out? Answer comes back the work is quietly worse and nobody says anything? You've built the same trap in miniature.

The nurses call the surgeon. That is the part that makes the system safe, and it's the part everyone forgets to build.

The thing you are depending on will be retired

One more, because it's the failure that catches people who have done everything else right.

On a day in June, a set of model identifiers we depended on were retired by the people who made them. Not deprecated with a long runway. Gone.

We had, by luck as much as foresight, a mapping layer that translated old names to current ones. The note recording it says that without it, every single call across the entire system would have broken.

Building your way of working around one tool from one company is an exposure, and that's what it looks like when it fires. They're not unreliable. The thing you tuned your habits to is a product, and products change underneath you without asking.

Using more than one isn't only about cost and it's not only about disagreement. And it's the reason a change at one vendor costs you an inconvenience rather than a week.

Deliberately, not by accumulation

One qualification before you go and do any of this, because there is a version of it that costs you money and buys nothing.

Use more than one deliberately, for reasons you have written down. Do not end up with more than one by accumulation, which is what happens when nobody is watching · a tool arrives for one job, another arrives for a second, and three years later there are five doing overlapping work, not one of them ever chosen against the others, and all of them being paid for every month.

That is all of the cost and none of the benefit, because the benefit was never in the number. It was in two of them being genuinely different, and in somebody having decided which was which and why.

The difference between those two situations is whether anybody is holding the whole picture. That is your job now, and it is the part that doesn't get cheaper.

What to do this week

Three things, none of which need any system.

Sort one week of your own work into judgement and typing. Be honest. Most of it's typing.

Put the typing on the cheap fast tool and see whether the output is worse. Usually it isn't, and the surprise is instructive.

Take one decision that actually matters and run it through two different tools. Read the gap. Not the answers, the gap.

That third one is where this chapter earns its place, and it takes about four minutes.

You will now have something no one else in your building has: a piece of work you can defend, produced for less, with the one contested paragraph already identified.

And every technique in the last five chapters has been about the work.

Not one of them has been about the person it's for. Which is the larger of the two problems, because a perfectly checked, well-scored, cheaply produced answer aimed at nobody in particular lands exactly the same way as no answer at all.

---

### ▪ DO THIS > > Sort one week, then run one decision through two tools. > > 1. Split last week's work into judgement and typing. Be honest. Most of it is typing. > > 2. Put the typing on the cheap fast option and see whether anything gets worse. Usually nothing does, and the surprise is the lesson. > > 3. Take one decision that actually matters — not a task, a decision. And put the same question into two tools running on different underlying models. Check that they do. Plenty of branded assistants share one engine. > > 4. Ignore both answers. Find where they part company. Write down that one point. > > 5. Spend your attention only there. > > Steps 3 to 5 take four minutes. Step 1 is a lunch break, once. > > When both answers agree: that's worth something, but it isn't proof. Two routes to the same place beats one route travelled twice, and it's still two routes. > > You'll know it worked when the disagreement points at the exact paragraph you'd have got wrong. And you find it in ninety seconds instead of in the meeting.

---

So before you send another thing, there is a room to read. And almost nothing that matters in that room was written down anywhere in the message you are answering.

Watch this chapter
A short illustrated film of Chapter 9. The narration is the author’s own words from the chapter. The animation, the voice and the score are AI-generated.
2,562 words