What this article covers
A participant at the talk asked a very good question. He had made a 150-page training deck at work, first using one model to check the content, structure and chapter order repeatedly, revising it more than ten times until he thought it was close. Then he had another model review it once, and it turned up a standards error in the content that the first model had missed even with deep thinking switched on. He later traced it to the difference between two international standards.
He did it more than ten times, both on the latest models, both with deep thinking on, and there was still a gap. Judging by this case, the problem looks more like the design of material, instructions and mechanism than something a stronger model would fix. Though one case cannot completely rule out model capability either.
My reading is this: when long documents go wrong, it is usually because three layers of design were never done, and most people only did the first, sometimes not even that. This article opens up all three, and gives you one passage of prompting that works better than "run it again".
- People who need an AI to check long documents: training materials, reports, standards, contracts, tenders
- People who have already switched models and turned on deep thinking and still do not feel confident
- People who want to turn a one-off check into a workflow they can reuse next time
- The order of the three checking layers and the criteria for each
- How two slicing methods cover each other, and why one is not enough
- A copy-ready prompt for discussing with the AI why it got something wrong
1. Why running it ten more times does not help
An uncomfortable fact first: with the same material, the same model and the same question, the tenth run usually does not differ from the second in any meaningful way. The wording will differ each time, and occasionally it will catch one or two extra things, but a blind spot at a given layer does not disappear because you ran it more times. There are three reasons, and they sit at different layers.
What the AI reads is not what you see. This is what data cleaning solves.
Read too much at once and it skips. This is what chunking solves.
The same model carries the same set of preferences. What it considered right, it still considers right on review. This is what switching vendors solves.
Running it again just circles inside the same layer. Ten runs later, the material is still that hard-to-read material, the slicing is still that swallow-it-whole slicing, and the perspective is still the same perspective.
2. Layer one, material: confirm what it reads matches what you see
There are already two articles on the site covering this layer thoroughly, so here are the links:
- Why Split Data into Cards? From PDF Readability to Card-Based AI Retrieval: why PDF and PPT are print formats made for humans while an AI reads coordinates and code, and how to break them up. For how a 150-page deck actually gets converted, read this one.
- How to organize data in the AI era and turn files into usable systems: what to do after cleaning, to organise it into a system both people and AI can work with.
So rather than repeat that, here are three points that bear directly on reviewing.
Serious standards material gets cleaned first
Training materials, regulations, standards, contracts: get one version number or one standard code wrong and the whole conclusion goes with it, and it goes wrong in a way that looks right. For this kind of content I always suggest converting to a plain-text Markdown file first, pulling the chapter structure and a content index out together, and then handing it to the AI. What comes out is a searchable copy for the AI to read, and the original has to be kept: tables, text inside images and footnotes are the things most likely to be lost in conversion, which is why the sampling check below goes back to the original to compare.
Do one sampling check after cleaning
Do not start reviewing the moment cleaning finishes. Pick three to five spots at random and compare them against the original: did a table skip a row, did a footnote get swallowed, did the text inside an image vanish entirely. Five minutes here saves ten wasted runs later.
Make "where it came from" part of the rule
This is what the participant's case suggested to me: since what was wrong was the standard version, then require it to cite sources in the checking rule.
What this does is turn things it says smoothly into things it must produce evidence for. Claims without a source surface by themselves.
But this does not stop the error in the case at the top. That participant's deck had two international standards mixed up. Both standards genuinely exist and both are citable, so when it cited the wrong one it could still show you a source, and it looked entirely reasonable. To stop that kind of error, the rule needs one more layer: specify which standard and which version this document is subject to, and then require it to compare every citation against that line.
The key to this layer is the third case. Most standards errors are not invented out of thin air. They are the wrong choice between two things that both hold, and when it chooses, it does not tell you it made a choice. What you have to do is turn "choosing" into an action it must report.
3. Layer two, slicing, and not just one way
The problem at layer two is that when it reads too much at once, it starts skimming.
Throw a hundred-plus pages in whole and the beginning reads clearly, but by the later parts it starts skipping. So chunk it. The real key is the next sentence: do not chunk it the same way twice.
Ten chapters means ten passes, one analysis per chapter, then one big consolidation of the ten results at the end.
Blind spot: cross-chapter contradictions are invisible, for example chapter two and chapter eight describing the same thing differently.
First have it group the ten chapters into five themes, then run one pass per theme.
Blind spot: ordering problems inside a chapter are invisible, for example a build-up from simple to complex that has been broken.
The two slicing methods cover different areas, and only overlapping them gives you coverage. This is what "use two or more different ways of chunking" actually means. It does not mean running the same slicing twice.
4. Layer three, perspective: have both vendors run once, then cross them
Layer three is where you actually switch models.
The method: run each slicing method through both vendors once, then cross the two sets of results. That leaves you with four sets of results, and what survives the crossing as disagreement is usually where a human genuinely needs to judge.
That participant's experience is the proof of this layer: both models on the latest version, both with deep thinking, more than ten repetitions, and still a gap.
My own addition: different models have different weightings, so they have different blind spots. This is not a question of which is stronger. It is that they care about different things.
5. More effective than one more run: discuss why it got it wrong
This is the step I think most people miss.
After finding an error, most people's reaction is "check it again". But the error has already happened, and running it again is just another roll of the dice. What works better is discussing the thing itself with it directly.
What I read out in class was this passage:
This passage does two things.
- It moves the problem up a layer: from "this one was wrong" to "we need to design a process that does not go wrong again". What you get back is not one correction. It is a checking design you can keep.
- It draws out limits it knows about but will not mention unless asked: which section it was not actually confident about, which comparison it cannot do, which content needs external data to verify.
6. Three sources of error: work out which one this is first
After the discussion you will need to decide which kind of problem this is, because all three are handled completely differently.
Sign: it is doing something different from what you wanted, but doing it earnestly.
Fix: write the criteria, scope and output format clearly.
Sign: it got this one right, but misses the same thing next time.
Fix: write the checkpoints, the order and the stop condition into a rule.
Sign: switching model, slicing and phrasing all produce the same error.
Fix: have a human review this part, or use external data that can be verified.
That participant's 150-page deck does not look like it exceeded the model's reading range, so I judge it as falling into the first two, a problem of instruction and mechanism. One thing to add: page count does not convert directly into how much a model can take in. What actually matters is the volume of text and the formatting after cleaning, so this is a judgement about this case, not a general rule.
7. How many rounds before you stop
Easier to miss than "how to review" is this: how many revisions before you stop.
If you only tell it "revise according to my standard" with no upper limit, it will keep revising until it thinks the result is good. That usually goes one of two ways: the bill keeps climbing, or the result drifts further off.
My work is mainly document analysis and reasoning, and my rule of thumb is two to three rounds. One round here means:
- Round one: A reviews B, and B reviews A
- Round two: A revises once following B's suggestions, B revises once following A's, and then they review each other again
There is a simple way to judge that it is time to stop: when the same problem sticks in the same place two or three times in a row, stop and report. Do not keep running.
After stopping, what you do is go back and look at what the rule left out. Usually it is one of these three:
- No definition of the criterion: you never said what "correct" means.
- No rule for exceptions: you never said what to do when a situation is not covered.
- No requirement for sources: you never asked it to show why something is right.
8. Last step: turn this into a process it runs on its own
At this point you have a complete checking design: cleaning and sampling, two slicing methods, two vendors crossing, discussing the mechanism, judging the three sources, stop conditions.
All of that is still "you walking it through each time". The last step is handing it over.
Judging the moment is simple: when you find this standard has run once or twice with no problems, you can tell it:
If it can, that is loop engineering. When it cannot, it is usually because it does not know what to do when something is wrong. So go back and add "how to notice it went wrong, how to fix it, when to stop" to the rule.
What that participant said at the end of the class stuck with me: once it is confirmed OK, turn it into a skill package. That order is right. Get the process running smoothly first, then save the smooth process, rather than trying to build a perfect skill package from the start.
How you can start
Take the document the AI most often misreads, pick three spots at random and compare them against the original. You will find out quickly whether the problem is in the material or the analysis.
Run the same document by chapter and by theme, and list the conflicts between them. Do not have it choose for you. The conflict list is itself the value.
Then open a new conversation, say only the sentence you would normally say, and test whether it starts up on its own.
None of the three steps needs programming. What you need is to state clearly what you check every single time, and that is something only you can do.
Three related articles
- Why Does the AI Keep Forgetting What I Told It?: From "I'll Remember That" to Letting It Run Itself: how to confirm the checking rule really was remembered and really gets triggered.
- How Do You Design Mutual Review Between Two AIs?: Three Levels of Review Mechanism: once layer three is in place, how deep the mutual review should go and how to word the prompt.
- Why Split Data into Cards? From PDF Readability to Card-Based AI Retrieval: the full method for layer one, how the material gets cleaned.