What is this article talking about?
AI writes your plans, proposals, and strategy documents fast and smoothly. But have you ever noticed that it goes along with you from start to finish? This article records how I had two AIs from different companies each write the same plan independently, then catch each other's mistakes, and how I kept the real decision for myself. I actually ran the whole thing once today, and the numbers and the points where it went wrong are all in the article.
- People already using AI to write proposals, plans, and reports, but with a nagging doubt about whether it's safe to hand it over as is
- People who have both Claude and ChatGPT and want to know where using the two together is strongest
- People submitting government grants, bids, or client proposals, who can't afford a single blind spot
- A complete five-step dual-track mutual-review process you can follow step by step
- Two copy-ready prompts, no tools to install
- A real-case failure list: AI acted out in advance how my proposal would die
If you're in a hurry, jump straight to "How you can get started"; the prompts are there.
Real scene
Today I am going to write a proposal for SBIR government R&D subsidies.
The starting point of the matter is that I saw a lot of companies on the market making AI CRM and enterprise information integration platforms, which automatically connect a lot of systems. My positioning is based on the division of labor with them: they are data plumbers, who collect structured data; I do knowledge extraction, turning unstructured documents into the organization's intellectual assets. This positioning should be written into a plan that can be submitted for review.
My old approach was to open one AI and revise back and forth with it. By the fifth round the document looked complete and every paragraph made sense. But then I noticed a problem: every round, it followed the direction I'd set in the previous round. When I said the division of labor was good, it laid out the division of labor persuasively; when I said nonprofits were the primary audience, it wrote about nonprofits with real feeling. The whole document was really just my own thinking amplified, with no one inside it playing devil's advocate.
With a document you submit for review, losing once means losing. The reviewers won't go easy on me.
Source of the problem
Broken down, a single AI writing important documents has three structural problems.
The same model, whatever logic is used when writing, will be used when reviewing. It cannot find holes in its own logic, just like when we proofread our own compositions, we will never find enough typos.
Once you state a direction, AI does its best within that direction. It rarely steps back to ask: is this direction even right? The sooner you hand it a conclusion, the sooner it stops questioning.
Important documents need to be challenged, but AI's default service posture is to get things finished. Between finished and correct lies a whole row of mines you can't see.
Some people say: just tell the same AI to switch roles and critique itself, wouldn't that fix it? I've tried. Switching roles changes what it calls itself, but not the underlying thinking habits. It still reviews itself with the same logic, and most of what it catches is surface-level.
Mechanism solution: planning dual-track mutual review loop
My solution is to let two models from different companies each run the whole process independently, then catch each other's mistakes. Today's setup is Claude on one track and Codex on the other (Codex is ChatGPT's desktop AI; the only thing that matters here is that it and Claude are brains trained by different companies). The whole process has five steps.
The finalized ideas, background documents, and official information are organized into the same package, and the two tracks start from the same starting line.
Each writes a first draft, has an advisor flag blind spots, has a mock review score it, and revises itself. The two processes can't see each other.
Exchange the finished work, you review mine and I review yours, against a fixed set of five review points.
Where they converge, adopt with confidence; where they diverge, list it as a decision point, and the AI is not allowed to decide on its own.
The integrated version goes back to the other company for one more review, to catch the integrator's own bias.
Step 0: Same input
I organized the finalized ideas, background files, and official information locations into the same package, and the two tracks received exactly the same input. Key discipline: The content transferred to the second track cannot carry any ideas already written in the first track. Both sides must start from the same starting line.
Step 1: Both tracks run independently, blind to each other
Each track does four things: write a first draft of the plan, flag blind spots from a strict advisor's perspective, run a mock target review that scores it, and revise itself based on the first two.
My advisors on this track are two persona skill packages, meaning a public figure's way of thinking organized into a role setting AI can play: Musk's first-principles perspective, and Naval's perspective on leverage and specific knowledge.
The Naval pass then pried the whole plan's center of gravity loose: rather than saying we'll distill knowledge, first prove that the quality of the distillation can be verified, build that measuring stick, and the four problems of quantification, R&D content, IP, and moat all disappear at once.
For the mock-review stage, I first had a sub-agent tasked with searching pull the official review materials out of my knowledge base: scoring dimensions, checkpoint format requirements, common point deductions, and eligibility red lines. Then I used that material to spin up three mock reviewers, one technical, one industry, one execution, each with a different picky streak. The three scored my first draft 2, 2, and 3 out of 5. There was only one shared reason it would fail: the whole plan didn't have a single qualifying quantitative checkpoint. The official text explicitly wants verifiable phrasing like "complete a given module, achieve a given percentage," and my first draft was all qualitative description.
At the same time, the Codex track ran to completion with no visibility into my track at all. It wrote its own version, flagged its own fifteen blind spots, had its own mock reviewers score it 57 out of 100, and listed eight possible reasons it could be rejected. It even caught a program-level issue: the official materials had no dedicated "local SBIR" page for Taipei City, so the proposal's eligible applicant category might need to change.
Step 2: Mutual review
Swap after both tracks are completed. You review my finished product, and I review your finished product. The focus of the review is fixed on five items: whether there is any misinterpretation of the original meaning, whether there are internal contradictions, whether the official information is quoted correctly, whether there are new blind spots missed by both parties, and whether the presentation of decision-making options is fair.
Step 3: Integrate, hand disagreements to a human
The track that initiated the run combines the two finished drafts and the two mutual-review notes into a single version. The iron rule: where the two tracks converge, adopt with confidence; where the two tracks diverge, the AI is not allowed to make the call itself, but must list it as a decision point and leave it to me.
Today's disagreement was who to target as the audience. One track argued for nonprofits, the other for small and micro enterprises. For this kind of business judgment, AI just gives the options and the pros and cons; making the call is my job.
Step 4: Final review by the other side
The integrated version came from my track, so I sent it to the other side for one last review. This pass caught the single most valuable cut of the whole run: when I presented the two audience options, I'd put my existing course income into the advantages column of one of them. The final review pointed out that course income proves execution ability, and using it as demand evidence that "someone wants to buy this service" is fooling myself. The demand evidence for both options is actually zero; they start from the same line.
What it guarantees and what it does not guarantee
Blind-spot coverage is far wider than a single model's, and where the two companies converge, the conclusion is highly credible; it guards against anchoring, because the second track starts from clean input; and the whole process leaves an auditable set of work files, so every change can be traced back to who raised it and at which pass.
The mock review is a rehearsal; it can't stand in for what the real reviewers actually think, so before you submit, still run it past a real person who has written or reviewed such cases. Everything AI produces is options and evidence; the responsibility for the final call can't be outsourced. Official materials go out of date, so re-check them against the original source before submitting.
How you can get started
You don’t need to install any tools, just open two different AIs and you can run the simplified version. Three steps.
The first step is to post the same requirement to the two AIs respectively. You can use this paragraph directly as the prompt:
The second step is exchange and mutual review. Paste A's final version to B, and B's final version to A:
The third step is to choose an AI to integrate the two versions and the mutual review opinions, and make your own decisions on any differences.
Advanced: Let the two AIs call each other directly
If you often use Claude and Codex at the same time, or want to run this loop semi-automatically, we recommend two open source packages:
- openai/codex-plugin-cc: lets you call Codex directly from inside Claude Code
- sendbird/cc-plugin-codex: lets Codex call Claude Code back the other way
My own division of labor: Claude handles analysis and reasoning, Codex handles execution. When the loop hits something undecided, the two models talk it out among themselves, and only call me in when they genuinely can't settle it. The smallest opening move for a knowledge worker is this: after you finish a piece of copy or a plan, have another model review it.
Ending Recap
- Three structural issues in writing important documents with a single AI: self-examination blindness, anchoring, and only seeking completion
- Dual-track mutual review loop five steps: same input, independent running, mutual review, integration to retain decision points, and final review against each other
- The live run: the mock review acted out the reasons it would fail in advance, and the final review caught a bias I couldn't see myself
- Trust what converges, send divergences to a human; simulation does not replace a real person
- With two prompts, you can open two windows and start running today
A reminder about my own positioning
Is this process slow? Slower than just opening one AI. But was speed ever what I actually wanted? What I want is for what I send out to withstand challenge, and for every run to leave the blind-spot list, the reviewer perspectives, and the decision record in my knowledge base, becoming the starting point for next time. Using AI to support decisions and accumulate knowledge assets is the point; speed is just a byproduct.
I'm Coach Jiang
Tacit knowledge distiller and AI application planner. I run two free online talks every month, sharing hands-on experience and methodology. If these topics interest you, if you want to keep learning, or if you have consulting needs, you're welcome to start with the community.
What I mostly cover: using AI as a thinking partner to raise the quality and depth of your decisions; and organizing knowledge and experience into prompts, skill packages, and knowledge bases so AI can apply them flexibly.
Welcome to join my LINE community, free lectures and methodologies are shared here first.
Join LINE community ↗