AI Application/AI Workflow

Let Two AIs Catch Each Other's Mistakes: My Plan Scored 2/5 From a Mock Review Before I Sent It

The dual-track mutual-review loop: two AIs from different companies each write the same plan independently, then catch each other's mistakes, and I keep the real decision for myself. A live record of the first run on a real government-grant proposal, with copy-ready prompts.

Dual-track loop cover: one terminal on each track, merging into a single document, with Mika holding the approval stamp
The two AIs each write a version and catch each other's mistakes.

What is this article talking about?

AI writes your plans, proposals, and strategy documents fast and smoothly. But have you ever noticed that it goes along with you from start to finish? This article records how I had two AIs from different companies each write the same plan independently, then catch each other's mistakes, and how I kept the real decision for myself. I actually ran the whole thing once today, and the numbers and the points where it went wrong are all in the article.

Who this is for
  • People already using AI to write proposals, plans, and reports, but with a nagging doubt about whether it's safe to hand it over as is
  • People who have both Claude and ChatGPT and want to know where using the two together is strongest
  • People submitting government grants, bids, or client proposals, who can't afford a single blind spot
What can you take with you?
  • A complete five-step dual-track mutual-review process you can follow step by step
  • Two copy-ready prompts, no tools to install
  • A real-case failure list: AI acted out in advance how my proposal would die

If you're in a hurry, jump straight to "How you can get started"; the prompts are there.

Real scene

Today I am going to write a proposal for SBIR government R&D subsidies.

The starting point of the matter is that I saw a lot of companies on the market making AI CRM and enterprise information integration platforms, which automatically connect a lot of systems. My positioning is based on the division of labor with them: they are data plumbers, who collect structured data; I do knowledge extraction, turning unstructured documents into the organization's intellectual assets. This positioning should be written into a plan that can be submitted for review.

My old approach was to open one AI and revise back and forth with it. By the fifth round the document looked complete and every paragraph made sense. But then I noticed a problem: every round, it followed the direction I'd set in the previous round. When I said the division of labor was good, it laid out the division of labor persuasively; when I said nonprofits were the primary audience, it wrote about nonprofits with real feeling. The whole document was really just my own thinking amplified, with no one inside it playing devil's advocate.

With a document you submit for review, losing once means losing. The reviewers won't go easy on me.

Source of the problem

Broken down, a single AI writing important documents has three structural problems.

1. Self-examination blindness

The same model, whatever logic is used when writing, will be used when reviewing. It cannot find holes in its own logic, just like when we proofread our own compositions, we will never find enough typos.

2. Anchoring

Once you state a direction, AI does its best within that direction. It rarely steps back to ask: is this direction even right? The sooner you hand it a conclusion, the sooner it stops questioning.

3. Just seek completion

Important documents need to be challenged, but AI's default service posture is to get things finished. Between finished and correct lies a whole row of mines you can't see.

Some people say: just tell the same AI to switch roles and critique itself, wouldn't that fix it? I've tried. Switching roles changes what it calls itself, but not the underlying thinking habits. It still reviews itself with the same logic, and most of what it catches is surface-level.

Mechanism solution: planning dual-track mutual review loop

My solution is to let two models from different companies each run the whole process independently, then catch each other's mistakes. Today's setup is Claude on one track and Codex on the other (Codex is ChatGPT's desktop AI; the only thing that matters here is that it and Claude are brains trained by different companies). The whole process has five steps.

STEP 0Same input

The finalized ideas, background documents, and official information are organized into the same package, and the two tracks start from the same starting line.

STEP 1Both tracks run independently

Each writes a first draft, has an advisor flag blind spots, has a mock review score it, and revises itself. The two processes can't see each other.

STEP 2Mutual review

Exchange the finished work, you review mine and I review yours, against a fixed set of five review points.

STEP 3Integrate; hand disagreements to a human

Where they converge, adopt with confidence; where they diverge, list it as a decision point, and the AI is not allowed to decide on its own.

STEP 4Final review by the other side

The integrated version goes back to the other company for one more review, to catch the integrator's own bias.

Dual-track loop five-step flow chart: same input, dual-track independent running, mutual review, integration of differences and submission, final peer review. After the fifth step, the dotted line circles back to the first step.
The five-step dual-track loop: after step five, it circles back to step one, forming a loop.

Step 0: Same input

I organized the finalized ideas, background files, and official information locations into the same package, and the two tracks received exactly the same input. Key discipline: The content transferred to the second track cannot carry any ideas already written in the first track. Both sides must start from the same starting line.

Step 1: Both tracks run independently, blind to each other

Each track does four things: write a first draft of the plan, flag blind spots from a strict advisor's perspective, run a mock target review that scores it, and revise itself based on the first two.

My advisors on this track are two persona skill packages, meaning a public figure's way of thinking organized into a role setting AI can play: Musk's first-principles perspective, and Naval's perspective on leverage and specific knowledge.

The harshest line from the Musk passIn the entire document, not one person has ever paid for the words "knowledge distillation." You have a lovely division-of-labor narrative, but narrative is just narrative; evidence is evidence.

The Naval pass then pried the whole plan's center of gravity loose: rather than saying we'll distill knowledge, first prove that the quality of the distillation can be verified, build that measuring stick, and the four problems of quantification, R&D content, IP, and moat all disappear at once.

For the mock-review stage, I first had a sub-agent tasked with searching pull the official review materials out of my knowledge base: scoring dimensions, checkpoint format requirements, common point deductions, and eligibility red lines. Then I used that material to spin up three mock reviewers, one technical, one industry, one execution, each with a different picky streak. The three scored my first draft 2, 2, and 3 out of 5. There was only one shared reason it would fail: the whole plan didn't have a single qualifying quantitative checkpoint. The official text explicitly wants verifiable phrasing like "complete a given module, achieve a given percentage," and my first draft was all qualitative description.

At the same time, the Codex track ran to completion with no visibility into my track at all. It wrote its own version, flagged its own fifteen blind spots, had its own mock reviewers score it 57 out of 100, and listed eight possible reasons it could be rejected. It even caught a program-level issue: the official materials had no dedicated "local SBIR" page for Taipei City, so the proposal's eligible applicant category might need to change.

Step 2: Mutual review

Swap after both tracks are completed. You review my finished product, and I review your finished product. The focus of the review is fixed on five items: whether there is any misinterpretation of the original meaning, whether there are internal contradictions, whether the official information is quoted correctly, whether there are new blind spots missed by both parties, and whether the presentation of decision-making options is fair.

Step 3: Integrate, hand disagreements to a human

The track that initiated the run combines the two finished drafts and the two mutual-review notes into a single version. The iron rule: where the two tracks converge, adopt with confidence; where the two tracks diverge, the AI is not allowed to make the call itself, but must list it as a decision point and leave it to me.

Today's disagreement was who to target as the audience. One track argued for nonprofits, the other for small and micro enterprises. For this kind of business judgment, AI just gives the options and the pros and cons; making the call is my job.

Step 4: Final review by the other side

The integrated version came from my track, so I sent it to the other side for one last review. This pass caught the single most valuable cut of the whole run: when I presented the two audience options, I'd put my existing course income into the advantages column of one of them. The final review pointed out that course income proves execution ability, and using it as demand evidence that "someone wants to buy this service" is fooling myself. The demand evidence for both options is actually zero; they start from the same line.

Why one AI can't catch itI was completely unaware of this bias. A single AI wouldn't catch it, because that bias grew out of my own narrative from the start.

What it guarantees and what it does not guarantee

Guaranteed What the dual track can do

Blind-spot coverage is far wider than a single model's, and where the two companies converge, the conclusion is highly credible; it guards against anchoring, because the second track starts from clean input; and the whole process leaves an auditable set of work files, so every change can be traced back to who raised it and at which pass.

Not guaranteed What the dual track cannot do

The mock review is a rehearsal; it can't stand in for what the real reviewers actually think, so before you submit, still run it past a real person who has written or reviewed such cases. Everything AI produces is options and evidence; the responsibility for the final call can't be outsourced. Official materials go out of date, so re-check them against the original source before submitting.

How you can get started

You don’t need to install any tools, just open two different AIs and you can run the simplified version. Three steps.

The first step is to post the same requirement to the two AIs respectively. You can use this paragraph directly as the prompt:

I want to write a [plan/proposal] with the following requirements: [Paste your needs and background]. Please complete the three steps independently: 1. Write a first draft. 2. Put on the role of a strict consultant and list 5 to 8 blind spots in this first draft, with a “how to fix them” for each. 3. Simulate [the target reviewer, such as the grant reviewer/client’s decision-making director], who rates the first draft, listing the tough questions he would ask and the most likely causes of death. Finally, revise it into your final version according to steps 2 and 3. Don’t ask me about the process, just run through it.

The second step is exchange and mutual review. Paste A's final version to B, and B's final version to A:

This is the same-topic plan another AI completed independently (below). Please only review it, do not rewrite it: 1. Does it misread my original intent? 2. Are there internal contradictions? 3. Are the facts and data correct? 4. New blind spots it missed that you can see. 5. Is the presentation of options biased?

The third step is to choose an AI to integrate the two versions and the mutual review opinions, and make your own decisions on any differences.

Advanced: Let the two AIs call each other directly

If you often use Claude and Codex at the same time, or want to run this loop semi-automatically, we recommend two open source packages:

My own division of labor: Claude handles analysis and reasoning, Codex handles execution. When the loop hits something undecided, the two models talk it out among themselves, and only call me in when they genuinely can't settle it. The smallest opening move for a knowledge worker is this: after you finish a piece of copy or a plan, have another model review it.

Claude and Codex dual-model workflow chart: analysis is handed over to Claude, execution is handed over to Codex, and the two suites allow both parties to call each other
Dual-model workflow: leave analysis to Claude and execution to Codex.
Two remindersFirst, order matters: let each side finish independently, then exchange. Exchange before writing and the second AI is already anchored, and the whole dual-track point is wasted. Second, use the convergence points with confidence and think through the divergences yourself; that is the one place in the whole process where the human cannot be absent.

Ending Recap

  • Three structural issues in writing important documents with a single AI: self-examination blindness, anchoring, and only seeking completion
  • Dual-track mutual review loop five steps: same input, independent running, mutual review, integration to retain decision points, and final review against each other
  • The live run: the mock review acted out the reasons it would fail in advance, and the final review caught a bias I couldn't see myself
  • Trust what converges, send divergences to a human; simulation does not replace a real person
  • With two prompts, you can open two windows and start running today

A reminder about my own positioning

Is this process slow? Slower than just opening one AI. But was speed ever what I actually wanted? What I want is for what I send out to withstand challenge, and for every run to leave the blind-spot list, the reviewer perspectives, and the decision record in my knowledge base, becoming the starting point for next time. Using AI to support decisions and accumulate knowledge assets is the point; speed is just a byproduct.

AI WorkflowCross-Family ReviewDecision SupportCase StudyFundamentals

I'm Coach Jiang

Tacit knowledge distiller and AI application planner. I run two free online talks every month, sharing hands-on experience and methodology. If these topics interest you, if you want to keep learning, or if you have consulting needs, you're welcome to start with the community.

What I mostly cover: using AI as a thinking partner to raise the quality and depth of your decisions; and organizing knowledge and experience into prompts, skill packages, and knowledge bases so AI can apply them flexibly.

Interested in AI × knowledge management?

Welcome to join my LINE community, free lectures and methodologies are shared here first.

Join LINE community ↗