Lesson 5 of 6 · Agentic Processes

Checking the work

A ten-minute research run hands you a report that looks finished. It is a draft from a new employee who is confident, fast, and occasionally makes things up. Here is how to review it in five minutes instead of trusting it in zero.

Everything in this module so far has built toward one moment: the report is on your screen, it has headings and numbered citations and a summary, and you have to decide what to do with it. Most people do one of two things. They read the summary and act on it, or they feel vaguely suspicious and ignore the whole thing. Both waste the run. This lesson is the third option.

The new employee rule

You already have a mental model for this. Imagine a sharp new hire, two weeks in, who you sent off to research something. They come back with a tidy write-up. You would not sign a contract on it without reading it. You would also not throw it out. You would read it with a specific question in mind: where is this most likely to be wrong, and what would it cost me if it were?

That is the whole method. A research run is a draft from a capable, fast, literal new employee who has no idea which of their sentences matter to you. Your review supplies the judgment they lack. The rest of this lesson is the specific habits that make the review fast.

Citations that do not exist, and citations that do

AI Basics lesson 9 covered hallucination in chat: the model predicts what a good answer sounds like, and sometimes that includes predicting what a citation looks like. Author, year, journal, page numbers, none of it real. In a plain chat, that is the main failure.

Research runs change the failure, not the risk. Because the run actually opened pages, most of its citations point to real pages. The new problems are subtler:

  • The page exists but says something else. The run read a page about national gutter guard prices and the report calls it a Southern Utah price. The link works. The claim does not.
  • The page exists but is weak. A forum post, a content farm, a vendor's own marketing, cited as if it were a study.
  • The synthesis invented the bridge. Source A says X, source B says Y, and the report says "therefore Z," which neither source said. The citation is on X and Y. Z has no source at all, and it is usually the most interesting sentence.
  • Precise numbers from nowhere. "61% of homeowners" with a citation to a page that has no survey on it. Round, confident, and fabricated in the write step.

The tell is the same one from lesson 9: names, dates, numbers, and anything with a percent sign. Those are where you look first.

A research summary. Tap each claim to check its source.

Checked 0 of 4

Micro-mesh guards are the most commonly recommended type for homes near desert dust and cottonwood fluff, because fine debris slides off instead of clogging. [1]

A 2024 Washington County survey found that 61% of homeowners with guards skipped gutter cleaning entirely the following year. [2]

Installed cost for micro-mesh runs $8 to $12 per foot in Southern Utah. [3]

Guards do not eliminate maintenance. Debris still collects on top and needs brushing off once or twice a year. [4]

Illustrative. The summary, the survey, the prices, and the sources are invented to show the three verdicts you will meet in real reports.

Sampling: check three, not thirty

You cannot open every citation in a forty-source report, and you do not need to. Sampling means checking a few deliberately chosen claims and using what you find to decide how much to trust the rest. Pick three:

  • The claim you will act on. The number going in the quote, the deadline going on the calendar, the vendor you are about to call.
  • The claim that surprised you. If it contradicts what you expected, either you learn something or you catch an error. Both are worth two minutes.
  • The most specific claim. The exact percentage, the exact dollar figure, the exact date. Specificity is where invention hides.

Open each one. Confirm the page exists, find the sentence the claim came from, and ask whether the report's version matches. Three for three, and you can read the rest with reasonable trust. One miss, and every number in the report gets opened before it leaves your desk. Two misses, and you rerun with a tighter brief.

Sampling feels lazy the first time. It is not. It is how auditors, inspectors, and editors work when the whole cannot be checked. The trick is that the three you choose are not random. They are the three where an error would hurt the most or hide the best, and that makes a small sample a strong signal about the rest.

Two prompts that make the report check itself

Before you spend your own minutes, spend the AI's. Two follow-up messages, sent right after the report lands, do a surprising amount of the work.

"Show me the source." Pick a claim and ask: "Quote the exact sentence from source 4 that supports the 90-day claim." If the run can produce the sentence, you read it and judge. If it comes back with a paraphrase, an apology, or a different source, you have found a soft spot without opening a single tab.

Ask for a confidence note. "For each claim in the summary, mark it high, medium, or low confidence, and for anything below high, say why." Models are imperfect judges of their own work, but they are decent at noticing where sources were thin or disagreed, because that happened during the run and is still in the context. The lows and mediums are your sampling list, handed to you.

Remember lesson 4: a chat checking its own draft is a weak checker. These two prompts are triage, not verification. They tell you where to look. Your eyes on the source page are still the check.

The five-minute review

Put together, here is the routine for any research or agent output you are going to rely on:

  • Source list first. Count them, skim the domains. Mostly one vendor, mostly forums, mostly older than your cutoff? Adjust your trust before reading a word of the body.
  • Ask for the confidence note. Get the report to flag its own weak spots.
  • Sample three. The claim you will act on, the one that surprised you, the most specific one. Open the pages.
  • Hunt the bridges. Find the sentences that draw a conclusion from two sources and ask whether either source actually drew it.
  • Decide by stakes. Low stakes, done. Medium, fix what you found and use it. High stakes, this report is the brief you hand to a professional, not the answer.

For a browsing agent, the same routine applies to the summary of what it did: open the confirmation email, check the total it quoted against the receipt, read the form it filled before the deadline passes. Actions leave a trail. Follow it the same way you follow a citation.

Try this yourself

Take the last research report you ran, or run the one from lesson 2, and send this as your first reply:

Before I read this, help me review it.

1. List every specific number, date, or percentage in the report, each with its citation number.
2. For each one, quote the exact sentence from the cited source that supports it. If you cannot find a supporting sentence, say "no direct support" instead of paraphrasing.
3. Mark each claim high, medium, or low confidence and give one line of reasoning for anything below high.
4. List any conclusion in the report that combines two sources in a way neither source states on its own.

Then open the three pages behind the lows. Time yourself. Most people finish in five minutes and find at least one thing they would not have caught reading top to bottom. That is the habit. It costs five minutes, and it is the difference between a research feature you can build a business decision on and one you cannot.

Next lesson6. Cost and time

Last updated August 24, 2026