AnyLearn
All lessons
Businessbeginner

Checking, When Checking Costs More Than Generating

Verification is now the expensive step, which changes what a sensible checking strategy looks like. This lesson covers the asymmetry between producing and refuting, deciding what to check before you read it, the questions that actually discriminate, and what a citation is worth.

Updated · AI-authored, review-gated · how lessons are made

Not signed in: your progress and quiz score won't be saved.
Progress1 / 8

The asymmetry

The economics of claims have inverted, and this is the fact that should reshape how anyone works.

Historically, producing a substantiated claim was expensive. Someone had to find the information, understand it, and write it down, and that cost limited how many claims entered circulation. Checking was comparatively cheap, because there were not many things to check and the effort of production filtered out most nonsense before it appeared.

That filter has gone. Producing a specific, confident, well-formed claim now costs almost nothing, and checking one costs exactly what it always did: finding the source, reading it, and confirming it says what was claimed.

Alberto Brandolini stated the general form of this as an aphorism, sometimes called the bullshit asymmetry principle: the effort needed to refute nonsense is an order of magnitude greater than the effort needed to produce it. That was already true, and the production side of the ratio has now fallen by a further large factor.

What follows practically, and it is uncomfortable.

Checking everything is not available. If a tool produces forty claims in a document and each takes four minutes to verify properly, verification costs more than writing the document by hand would have. Any policy of check everything will be abandoned within a week, and abandoned policies are worse than honest ones because people believe the checking is happening.

So the realistic strategy is not more checking. It is deciding, in advance and by category, what gets checked and what does not, and being honest about the rest.

That decision is the subject of this lesson, and it has to be made before the output arrives, because the previous lesson established that judgement at the moment of reading is compromised.

Full lesson text

All 8 steps on one page, for reading, reference, and search.

Show

1. The asymmetry

The economics of claims have inverted, and this is the fact that should reshape how anyone works.

Historically, producing a substantiated claim was expensive. Someone had to find the information, understand it, and write it down, and that cost limited how many claims entered circulation. Checking was comparatively cheap, because there were not many things to check and the effort of production filtered out most nonsense before it appeared.

That filter has gone. Producing a specific, confident, well-formed claim now costs almost nothing, and checking one costs exactly what it always did: finding the source, reading it, and confirming it says what was claimed.

Alberto Brandolini stated the general form of this as an aphorism, sometimes called the bullshit asymmetry principle: the effort needed to refute nonsense is an order of magnitude greater than the effort needed to produce it. That was already true, and the production side of the ratio has now fallen by a further large factor.

What follows practically, and it is uncomfortable.

Checking everything is not available. If a tool produces forty claims in a document and each takes four minutes to verify properly, verification costs more than writing the document by hand would have. Any policy of check everything will be abandoned within a week, and abandoned policies are worse than honest ones because people believe the checking is happening.

So the realistic strategy is not more checking. It is deciding, in advance and by category, what gets checked and what does not, and being honest about the rest.

That decision is the subject of this lesson, and it has to be made before the output arrives, because the previous lesson established that judgement at the moment of reading is compromised.

2. Triage by consequence, not by suspicion

The instinct is to check the claims that seem doubtful. That is precisely the wrong criterion, because the previous lesson established that your sense of doubtfulness has been disabled.

A generated claim that is wrong does not feel different from one that is right. Selecting by suspicion therefore samples essentially at random, while feeling deliberate.

The criterion that works is consequence. What happens if this particular claim is false?

That gives a workable triage.

Always check, regardless of how confident it looks. Anything a decision rests on. Anything going to a customer, a regulator, a court or a funder. Every number. Every citation, name, date and quotation. Anything about a specific person. Anything legal, medical, financial or safety-related. And anything you will be personally associated with.

Spot check. Background and context, where a single error would be embarrassing but not consequential. Sample a few claims; if the sample is clean, the rest is probably reasonable, and if the sample contains an error, check everything because the output is unreliable for this topic.

Do not check, honestly and deliberately. Framing, structure, phrasing, brainstormed options you will evaluate anyway, and anything where you are using the output as a prompt for your own thinking rather than as information.

Why stating that third category explicitly matters. A person who has decided that this category is unchecked will not later mistake it for verified. The failure mode is not using unchecked material; it is losing track of which material was checked.

And the useful property of consequence-based triage is that the list is short. Numbers, citations, names, and anything that reaches a decision or a third party. Most output is none of those.

3. What checking actually means

The word covers several different activities with very different value, and confusing them is why people believe they have verified something when they have not.

At the bottom, and worth naming as worthless: asking the model whether it is sure. A system that produced the claim is not an independent check on it, and sycophancy means pushing back frequently produces agreement in either direction.

Also near-worthless: asking a second model. Different systems trained on overlapping material make correlated errors, so agreement between them is weak evidence. It feels like corroboration and is closer to asking the same question twice.

Weak: checking that the claim is consistent with what you already believe. This catches claims that contradict your knowledge and is blind to everything in the space where your knowledge is thin, which is exactly where you were relying on the tool.

Better: finding the claim asserted somewhere else independently. Real evidence, with the caveat that the web now contains a great deal of generated content, so agreement across sources is weaker corroboration than it used to be.

And the real thing: going to the primary source and confirming it says what was claimed. Reading the paper, the statute, the filing, the documentation. This is the only step that reliably discriminates, and it is the one that costs four minutes.

The practical implication of that hierarchy. If a claim is in the always-check category, only the last row counts. Everything above it produces the feeling of verification at a fraction of the cost, which is precisely why people do it.

flowchart TD
A["Ask the model if it is sure"] --> B["Worthless: not independent, and sycophancy answers either way"]
C["Ask a second model"] --> D["Near-worthless: correlated errors, feels like corroboration"]
E["Check against what you already believe"] --> F["Weak: blind exactly where your knowledge is thin"]
G["Find it asserted independently elsewhere"] --> H["Real, but weakened by generated content on the web"]
I["Read the primary source"] --> J["The only step that reliably discriminates"]

4. Citations, and the three ways they fail

Citations deserve specific treatment because they are the element people most often accept as evidence of checking, and they fail in three distinct ways that require different responses.

The reference does not exist. A plausible author, a plausible title, a plausible journal, a plausible year, referring to nothing. This is the widely-publicised failure, it has produced sanctions against lawyers who filed briefs containing invented cases, and it is the easiest to catch: search for it and it is not there.

The reference exists and does not say what was claimed. Substantially more common and far harder to catch, because the citation checks out at the level most people check. The paper is real, the authors are real, and the finding attributed to it is not in it, or is a weaker version, or is something the paper explicitly argued against. Catching this requires reading the source, not confirming it exists.

The reference exists, says roughly that, and is a poor support. A preprint that failed replication. A study with a sample of twelve. A finding since superseded. A source misrepresenting the field's actual state of agreement. This requires domain judgement, and it is where a citation can be technically defensible and still misleading.

What follows. Confirming a citation exists is not checking it, and it is what most people do. The second failure mode is more common than the first, so a workflow that only catches non-existent references will pass most of the problems through.

And the general rule for anything with a source attached: if you have not read it, do not cite it. That was always the standard in careful work, and the only thing that changed is how easy it became to violate.

5. Questions that discriminate

When you cannot verify against a source, some questions still separate reliable output from unreliable, because they probe where these systems are weak rather than where they are strong.

Is this the kind of thing that would be written down? Models represent well what appears often in text. A claim about a widely-documented topic is more likely reliable than one about something niche, recent, proprietary, or local. Your company's internal process, a small jurisdiction's rules, last month's release: all are places where output should be assumed unreliable.

Does this depend on a specific version, date or jurisdiction? Anything with an edition, a version number, a legal jurisdiction or a recent change is high-risk, because the correct answer varies along a dimension the model handles poorly. Software APIs, regulations, tax rules, medical guidance.

Is this suspiciously well-suited to what I asked for? A model asked for evidence supporting a position will supply it. Output that fits your framing perfectly, with no tension and no inconvenient exception, is a signal about the question rather than the world.

Would the opposite claim have been produced as fluently? A useful thought experiment. If you can imagine asking the inverse question and receiving an equally confident answer, the confidence is a property of the format rather than the evidence.

Are the specifics load-bearing? Vague claims are relatively safe. Precise ones, a percentage, a date, a section number, a named individual, are where fabrication concentrates, because specificity is generated as readily as anything else and cannot be produced reliably from a statistical pattern.

And the summary heuristic that captures most of these. The more specific, recent, local or niche a claim is, the less you should rely on it without checking, and those are exactly the properties that make a claim useful.

6. Asking the question the right way round

A great deal of the reliability problem is determined before any output exists, by how the request was framed.

The framing that causes trouble. Give me evidence that remote work improves productivity. This asks for one side, and it will be supplied, complete with studies and figures. The output is not lying; it is answering the question, and the question was for advocacy.

People then read the result as though it were a survey of the evidence, because it looks like one.

Better framings, which change the output substantially.

What does the evidence say about remote work and productivity, including where it conflicts. This asks for the state of a question rather than for a position.

What is the strongest case against this, made by someone who disagrees. This produces the objections you will face, and it is the single most useful prompt in this lesson.

Where is this contested, and by whom. Genuine disagreement is usually visible in the literature, and asking for it surfaces the fact that a question is open, which a one-sided answer conceals.

What would have to be true for this to be wrong. Harder to answer agreeably, which is why it works.

And a specific discipline from the previous lesson, worth repeating because it is the most violated. Do not state your view before asking. Once your position is in the conversation, sycophancy has something to agree with, and every subsequent response is contaminated.

The general principle. You are not extracting facts from a database. You are getting a response shaped by your request, and a request shaped like an argument returns an argument. Most complaints about unreliable output are complaints about advocacy that was requested without anyone noticing.

7. Keeping track of what was checked

A practical problem that causes real damage and receives almost no attention: within a working session, checked and unchecked material become indistinguishable.

How it happens. You generate a draft. You verify three of its claims because they mattered. You edit, restructure, and combine it with other material. A week later you use it in something else. At that point nothing on the page distinguishes the verified claims from the fifteen you never touched, and your memory of which was which has gone.

The result is that unchecked material acquires the status of checked material through nothing but the passage of time and reuse. This is how a plausible invented statistic ends up in a board paper eighteen months later, having been repeated internally until it became a thing everyone knows.

What prevents it, proportionate to the work.

Mark claims as you verify them, in the document, while you are doing it. A source link next to a number is the whole mechanism, and it survives editing in a way that memory does not.

Keep unverified material visually distinct while drafting. Brackets, a colour, a note. Anything that makes the unchecked status visible rather than remembered.

Before anything leaves your hands, sweep for unmarked specifics: numbers, dates, names, citations. If a number has no source next to it, either find one or remove it.

And for anything that will be reused, record where the material came from. Not for provenance in the abstract, but because the person reusing it in six months is you, and you will have no way to tell.

The underlying observation. Verification does not persist in the artefact unless you put it there. An unmarked checked claim and an unmarked invented one are the same object.

8. The checking policy

Everything above, as something you can actually operate.

Decide the categories in advance, not while reading. Judgement at the moment of reading is compromised, which the previous lesson established, so the decision has to be made when nothing is in front of you.

Always verify against a primary source: every number, every citation, every name and date, anything about a specific person, anything a decision rests on, anything reaching a customer, regulator or court, and anything with your name on it.

Spot check background and context. If the sample contains an error, check everything on that topic, because it tells you the output is unreliable there.

Deliberately do not check framing, structure, phrasing and options you will evaluate anyway. Say so, so it is never mistaken for verified.

Remember what checking means. Asking the model again is worthless. Asking another model is nearly worthless. Consistency with your own beliefs is blind where your knowledge is thin. Only the primary source discriminates, and for citations that means reading it rather than confirming it exists, because saying-something-different is more common than not-existing.

Frame requests as questions rather than as arguments, ask for the strongest case against, and never state your view first.

Raise suspicion for anything specific, recent, local, niche, or dependent on a version or jurisdiction. Those are exactly the useful claims and exactly the unreliable ones.

And mark what you verified, in the artefact, as you go, because verification does not survive editing and reuse unless it is written down.

The honest summary. This is not a higher standard than careful people always applied. It is the same standard, made harder to apply by the volume, and made necessary by the removal of the cost that used to filter claims before they reached you.

Check your understanding

The lesson ends with a 5-question quiz. Take it in the player above to see your score.

  1. Why is 'check everything' not a workable policy?
    • Verification is unreliable
    • Producing claims is now nearly free while checking costs what it always did, so verification can exceed the cost of doing the work by hand
    • Most claims cannot be traced to a source
    • It duplicates the model's own checking
  2. Why is suspicion the wrong criterion for deciding what to check?
    • Suspicious claims are usually correct
    • It takes too long to form an impression
    • A wrong generated claim does not feel different from a right one, so selecting by suspicion samples at random while feeling deliberate
    • Suspicion is influenced by the topic rather than the claim
  3. Why is asking a second model near-worthless as verification?
    • Second models are usually less capable
    • They cannot access the original prompt
    • It doubles the cost for no gain in speed
    • Systems trained on overlapping material make correlated errors, so agreement is weak evidence
  4. Which citation failure is most common and hardest to catch?
    • The reference exists but does not say what was claimed
    • The reference does not exist at all
    • The reference is formatted incorrectly
    • The reference is behind a paywall
  5. Which properties make a claim least reliable?
    • Length and technical vocabulary
    • Being specific, recent, local, niche, or dependent on a version or jurisdiction
    • Being expressed with hedging
    • Appearing without a citation

Related lessons

Business
beginner

Why Fluent Text Defeats Your Judgement

Generated output is persuasive through properties unrelated to whether it is true. This lesson covers processing fluency, automation bias, the illusion of explanatory depth, and sycophancy: four mechanisms that make a confident draft harder to evaluate than a hesitant colleague.

8 steps·~12 min
Business
beginner

What Happens to Your Own Thinking

Delegating cognitive work has effects on the delegator. This lesson covers cognitive offloading and what is known about it, why the tasks that feel wasteful are often where skill is built, the expertise paradox in who benefits, and how to decide what to keep doing yourself.

8 steps·~12 min
AI
intermediate

What This Teaches About Measuring Anything

The exchange is a case study with transferable rules. A conclusion resting on failures needs a failure taxonomy. Every instance must be verified solvable before anyone is scored against it. Output format is a confound whenever answers get long. And when two explanations fit the same data, the productive move is to find the prediction on which they differ, then test it.

8 steps·~12 min
AI
intermediate

The Rebuttal: Three Ways to Score Zero Without Failing

The response disputed none of the data and argued the experiment measured something other than reasoning. Models had to print move lists exceeding their output limits, and said so in the transcripts. Some instances had no solution and were scored as failures anyway. And asking for a program instead of a move list produced high accuracy on instances reported as total collapse.

8 steps·~12 min