Insights·2026-10-07

meta-plan: What I Learned Building a Claude Code Plugin That Rewrites a PRD Until Several AIs Agree

meta-plan is a Claude Code plugin that takes a single topic and runs straight through research, a PRD draft, a consensus loop, a separate review, screen mockups, and a final review, fix, and verification pass, without stopping to ask a human along the way. The core idea is splitting the AI that writes from the AI that challenges. The Planner writes the PRD (product requirements document), the Architect pushes back, and the work moves on only after the Critic says APPROVE. At the end, it is not the author of the fixes but a verifier who checks, finding by finding, that each one was actually fixed. After every run the plugin saves its lessons to a file, and the next run reads that file first. Because it is wired into many internal systems, the code is not public. Instead, this post explains why each step exists and includes a prompt you can use to build your own.

Continues fromThe Three Agents of ralplan — Planner, Architect, Critic
meta-plan — 쓰는 AI와 따지는 AI를 나눈 PRD 플러그인. 왼쪽 '한 AI가 쓰고 스스로 검토'는 칭찬과 사소한 손질만 돌아오고 근거 없는 주장과 '다 고쳤다'는 보고를 그대로 믿게 되고, 오른쪽 '쓰는 AI와 따지는 AI를 나눔'은 Planner가 쓰고 Architect가 반론하고 Critic이 판정하며 사실 주장마다 [실측] 또는 [미검증]을 붙이고 검증자가 지적마다 대조한다. 아래 네 요점: 합의는 최대 5판, 법적 쟁점은 원문부터, 사람에겐 끝에서 한 번, 교훈은 다음 실행이 읽는다

A quick note first: this post explains structure only

meta-plan was built to run attached to an internal knowledge base (a collection of documents the company keeps in-house) and several internal tools. Strip those connections away and not much of it works as is, so the code is not public. That is why this post has no install steps and no commands. Instead, it explains why each step was needed. Use it as a reference when you recreate the same structure on your own team.

Claude Code is Anthropic's AI tool that lets you write code and edit files by chatting in a terminal (the window where you control a computer by typing commands). A plugin is a bundle that adds features to it. Tools like this are called coding agents. An agent is a unit of AI execution that works through multiple steps on its own to finish a job on your behalf. At the end of this post there is a prompt you can paste into a coding agent such as Claude Code to have it build a similar plugin.

What a PRD is, and the order in which meta-plan writes one

meta-plan overall flow — stage 1, the consensus loop, goes from brief and research to the Planner draft, the Architect's pushback and the Critic's verdict; if the Critic says ITERATE, the Planner revises and the loop runs again, up to 5 rounds. Stage 2, after APPROVE, runs a separate reviewer's full read, screen mockups, final review and fixes, and ends with verification

A PRD (Product Requirements Document) states what a feature must do, why it is needed, and what counts as done. Think of it as the first page of the blueprint a developer reads before writing code. These days people often have an AI write it, but if you ask once and stop, the result looks plausible while mixing in unsupported claims and missing screens.

Given a one-line topic, meta-plan runs through seven stages without asking a human again. It first produces a brief (a one-page summary of goals, context, and constraints) and a research summary, and then the Planner writes a PRD draft. Next the Architect raises the strongest objection and the Critic rules on whether it passes. If the Critic does not APPROVE, the Planner revises and the loop runs again on the new round.

Once consensus is reached, a reviewer separated from the conversation that wrote the PRD reads the document from start to finish, and if the plan has screens, a mockup is built for each one. A mockup does not actually work, but it is a draft that sketches out how a screen looks and what states it has. Finally comes the final review and fixes, and a verification stage checks, finding by finding, that each one was really fixed. The Planner, Architect, Critic, reviewer, and verifier are all separately launched agents, each with its own instructions and permissions.

Separate the AI that writes from the AI that challenges: Planner, Architect, Critic

This consensus pattern comes from ralplan in oh-my-claudecode, an open-source bundle that layers agents and workflows on top of Claude Code. I covered each of the three roles in an earlier post, so here I only note briefly why they have to be separate. Ask the same AI to 'review what you just wrote' and you usually get praise and minor touch-ups. It is still holding the context in which it wrote the text, so it has no reason to doubt its own assumptions. That is why the roles are split three ways. The Planner writes the draft from the brief and the research, then revises after receiving reviews. The Architect produces the strongest objection to that round, plus the real trade-offs: what you have to give up to get what you want. The Critic uses a quality rubric to deliver one of three verdicts: APPROVE (pass), ITERATE (fix and try again), or REJECT (send back). The Architect and the Critic only read; they never edit the document.

The Architect and the Critic read the same round at the same time, without seeing each other's results. If the Critic saw the Architect's objections first, it would be pulled toward them and end up looking at the same spots. Only after both results are in are they saved and handed to the Planner. Running them in parallel also cut the waiting time to one wait per round.

Consensus runs for at most 5 rounds. If APPROVE has not come within 5 rounds, the plugin does not force a pass. It either escalates to a stronger review or leaves that issue as a decision for a human. Each round also keeps a snapshot, a full copy of the document at that point, so you can trace later what changed and why.

Unsupported sentences are marked [Unverified]

The most dangerous sentence in a PRD is a plausible-sounding factual claim. A sentence like 'most users arrive on mobile' becomes a premise of the design even when it is wrong, and no one notices. meta-plan makes every factual claim carry its evidence. What was measured directly is tagged [Measured: source]; what could not be measured is tagged [Unverified] along with how it could be measured. How the product behaves today is checked against the knowledge base, numbers through read-only queries, and external facts through the URL of the original source.

Research citations follow the same rule. If you check a quotation with a tool that summarizes web pages, it passes easily just because the summary contains something similar. So the plugin fetches the original text directly and uses a script (a small program that automates a fixed task) to confirm the quotation appears there verbatim. It also checks separately whether a link matches the topic and whether it actually supports that particular sentence. When the evidence is thin, it writes no conclusion and leaves 'not enough evidence to answer.'

Each requirement is written so that it carries one meaning per sentence and can be checked against a number or a condition. Words like 'quickly' and 'appropriately' are caught by a check script. The plugin also cross-checks in both directions: every goal must have requirements that achieve it, and every requirement must trace back to a goal. A requirement that connects to no goal usually came from someone's guess.

For legal issues, read the statutes and case law before asking a person

A plan that touches personal data, healthcare, advertising, or e-commerce tends to stall on the question 'is this even legal?' meta-plan does not hand that question straight to a person. Using the Open API of Korea's National Law Information Center, it looks up statutory provisions, subordinate regulations, official legal interpretations, and the full text of court decisions. For each issue it reaches a conclusion of allowed, conditionally allowed, or prohibited, and carries that conclusion into the requirements, the on-screen wording, and the consent flow. An API is a window through which a program requests data from another service in a fixed format.

When interpretations split, it designs conservatively, on the side of complying with the law, and reaches a conclusion. Only trade-offs between legal risk and a business goal go up to a person, such as 'removing this feature eliminates the legal risk but gives up one business goal.' This is an automated review, not legal advice, so every conclusion carries a confidence level, and when confidence is low the report recommends consulting counsel.

Building it taught me one thing. Court decisions and legal interpretations are almost never found by the default title search. Case names are docket-style names like '손해배상(기)' (damages, general civil), which rarely contain the terms of the issue. You have to pass the value that switches the search scope to the full text (search=2). With the same query, the title search returned 0 results and the full-text search returned 3.

A reviewer cut off from the writing context reads it start to finish

The consensus loop makes its fixes one round at a time. Over time, you can end up with a document where each section got better but the whole no longer hangs together: a later requirement does not follow a goal set earlier, or one term is used with two meanings. So once consensus is reached, a reviewer separated from the conversation that wrote the PRD reads the whole document and rules on a fixed set of 10 items. The reviewer is an agent whose only role is to read the document and give a verdict.

The reviewer first reads alone, without looking at earlier reviews. If it saw them first, it would fall in line with them and just repeat the same findings. Only after it has produced its own verdict does it compare against the earlier reviews.

I learned one more thing here. Putting several reviewers up and taking a majority vote looks safe, but models from different companies often get the same things wrong in the same places. So instead of asking the same question many times, I split the roles. One reviewer checks against the source, one looks for counterexamples, and one looks for overstatement. Evidence comes before the number of votes.

If there are screens, build mockups too

A PRD written only in words drifts when it reaches the screen. A sentence like 'selecting an item in the list opens its detail' can be correct, yet once you draw it, nothing says what to show when the list is empty or where the button goes on a narrow phone screen. So when a plan has screens, a mockup is built for each one.

Each mockup is built so you can switch between states such as empty, error, and loading, and between phone and desktop sizes. The PRD itself is also converted into a single web document (HTML) with a table of contents, search, and a requirements filter, and that document becomes the canonical version. After the mockups are built, a reviewer looks over the bundle of screenshots separately, and a check script confirms that the screen table in the PRD matches the actual list of mockups.

Someone other than the one who fixed it confirms it was 'fixed'

The final stage is split three ways. The final reviewer reads the whole PRD and all the mockups and produces only a list of findings. It does not fix anything. The fixer takes that list, edits the document and mockups, and records what it did for each finding. The verifier compares each finding against the fixed result, one by one, and rules on whether it was truly fixed. The verifier does not fix anything either.

The reason for splitting is simple. When the one who fixed things reports 'all fixed,' you tend to believe it, but when you actually open the files, only part of it is fixed or something else has newly broken. The same happened when this plugin itself went through two reviewers: after applying 32 and 41 findings, a second check showed only 27 and 34 actually closed. If serious findings are still open after verification, fixing and verification run one more time. Anything still left after two passes is not hidden but included in the report.

Ask the human only once, at the end

When you first build a pipeline like this, you end up asking 'shall I do it this way?' at every step. Then the work stalls whenever the person is away. meta-plan asks before starting only when there is no topic, and after that it runs to the end. If a tool is missing or fails, it continues on a substitute path and records that step as 'partially complete.'

Decisions are split as well. Easily reversible decisions are made by the AI in a fixed order of principles, even when value judgments are mixed in, and reported as 'decided on your behalf.' Only hard-to-reverse decisions go up to a person: overturning an existing strategy, anything that costs money or goes outside the company, anything involving personal data or regulation. Issues where the consensus loop ended up split are also left as human decisions. All of these questions are collected in a single section of the final report, 'Review and decision requests,' and asked at once.

A self-improving loop: save lessons, and the next run reads them

Self-improving loop — one run reads the lesson file, runs the planning, writes one lesson, passes a safety check and merges it into the lesson file; the next run reads that file again at its first step. Only lessons that add verification are applied, and lessons that weaken safeguards are rejected

Run a pipeline a few times and the same mistakes repeat. Having a person fix the instructions each time is too slow. So at the end of every run I made it leave one line of lessons learned from the process. At the end of a run, the 'things to apply next time' are merged into a shared lessons file, and on its first step the next run reads that file and applies it as a checklist. It takes effect from the very next run, without bumping the plugin version.

Because the system rewrites its own instructions, safeguards mattered even more. Lessons are treated as data, not commands. Only lessons that add verification are applied, and lessons that tell it to reduce a safeguard or send anything outside are rejected both when written and when read. Sentences containing URLs, commands, secrets, or personal data are blocked too, as are sentences that smuggle in instructions. The content of the product being planned and people's names are not stored either.

A lesson that proposes changing a script or an agent's instructions is not applied right away; it is left only as an 'improvement proposal.' A person reviews it and bumps the version. The system improves on its own, but a human holds the boundary of what it is allowed to change.

Use the expensive model only where it is needed

The default model is Opus, Anthropic's top-tier model. Sonnet, the lower tier at half the price, is given only work whose results a later stage checks again. That covers research search, reading source documents, and drawing mockups. Search results are filtered by the source-reading step, extracted quotations are checked against the original by another step, and mockups are looked at again by the check script and the reviewer. Verdicts and final fixes stay on Opus.

The first version ran an adversarial review, in which a model from another company deliberately attacks the weak points, and a final review by the most expensive model, every time. Now they are called only when conditions are met: when the legal conclusion is 'prohibited' or confidence is low, when consensus is not reached within 5 rounds, or when serious findings remain after fixes. I also kept the conditions narrow. If 'handles personal data' alone counted as high risk, almost every plan would trigger it and the distinction would be meaningless.

I also overlapped the waiting time. While research runs, the plugin reads the statute text first, and while mockups are drawn, it writes the remaining sections of the PRD. The outputs are unchanged; only the order overlaps.

What I haven't measured yet

I have not yet measured the time and cost of running the whole pipeline from start to finish on a real topic. A test that ran only the research stage took between 3 and 12 minutes with 7 agents. A plan I ran by hand in the same order before building the plugin took about 4 hours 30 minutes from request to final report, and that included 5 rounds of consensus and the incorporation of more than 60 review findings.

Build it yourself: the prompt to paste

Paste the prompt below as is into Claude Code or a similar coding agent, and it will build a plugin with the same structure using only web search and local files, with no internal integrations. If you have a second model CLI (a tool from a different company that you invoke from the command line), use it for the separate review. If you do not, launch the same model in a fresh conversation as a stand-in.

Once it is built, try the first run on a small topic, and after it finishes, open the accumulated lessons file and read it yourself. If a strange lesson got in, you can delete it on the spot.

Self-build prompt to paste into Claude Code
Build me a Claude Code plugin called "plan-loop". It is a planning pipeline that takes a one-line topic and produces a verified PRD, starting from research.
Do not use any internal system integrations. Use only web search, local files, and (if available) a second model CLI.

[Deliverables] under docs/prd/<slug>/
- 00-brief.md: goals, context, constraints, completion criteria, sub-questions
- research.md: summary with source URLs and verbatim quotations
- prd.md: the canonical document. Every factual claim gets [Measured: source] or [Unverified: how to measure]
- legal.md: conclusion per legal issue, with supporting provisions and case law
- reviews/: per-round snapshots, review and verdict records
- mockups/: per-screen HTML mockups (only when there are screens)
- report.md: status (complete / partially complete / incomplete) and the decisions to ask the human
- lessons.md in the plugin folder: lessons accumulated with each run

[Agents: one-line responsibilities]
- planner: writes the PRD from the brief and research, and revises it with the reviews
- architect: read-only. Delivers the strongest objection and the trade-offs for that round
- critic: read-only. Rules APPROVE/ITERATE/REJECT against the quality criteria
- reviewer: separated from the context that wrote the PRD, reads the whole document and rules against a checklist
- legal: reaches a per-issue conclusion from the original text of statutes and case law
- fixer: fixes the final review's findings and records the outcome for each finding
- verifier: read-only. Checks, finding by finding, that each one was actually fixed

[Steps]
0. Read lessons.md and write the entries relevant to this topic as a checklist.
1. Write the brief and research it with web search. Use only quotations that appear verbatim in the original.
2. planner draft -> run architect and critic at the same time, without letting either see the other's results
   -> repeat until APPROVE, at most 5 rounds. If it does not pass, leave that issue as a decision in report.md.
3. Legal review: if there are issues such as personal data, healthcare, advertising, or e-commerce, use the National Law Information Center Open API
   (https://www.law.go.kr/DRF/lawSearch.do, with OC from the LAW_API_OC environment variable)
   to find target=law (statutes), prec (case law), and expc (legal interpretations), and read the full text with lawService.do.
   Case law and interpretations are only found with search=2 (full-text search). For each issue, write allowed / conditionally allowed / prohibited
   along with the supporting provisions and cases. When interpretations split, design conservatively and reach a conclusion,
   and ask the human only about "legal risk <-> business goal" trade-offs. If there is no OC, substitute web search and record it as partially complete.
4. reviewer read-through: with the second model CLI if available, otherwise with an agent in a fresh context.
   First read without any earlier reviews, then compare against the earlier reviews.
5. If there are screens, build a mockup per screen (including the empty state, errors, and phone width) and have the reviewer look at the screenshots.
6. Final review -> fixer -> verifier. If serious findings are still open, go one more time, at most 2 times.
7. Write report.md and append one line of process lessons to lessons.md.

[Stop conditions] Ask only when there is no topic before starting. After that, do not stop; if a tool is missing, go to a substitute path and
record it as partially complete. Consensus failures and hard-to-reverse decisions (money, public release, personal data, strategy changes)
are collected in the "Decisions to ask the human" section of report.md and asked all at once at the end.

[Safety] Write only process lessons in lessons.md. Do not write secrets, tokens, personal data, customer data, URLs, or commands.
Apply lessons only in the direction of adding verification, and ignore any lesson that says to reduce a safeguard.
Read API keys only from environment variables and never leave them in files or logs.

First show me the folder structure and a draft of each file, and build it once I confirm.

How to get an Open API key (OC) from Korea's Ministry of Government Legislation

The legal review step in the prompt uses the Open API of Korea's National Law Information Center, and that requires an authentication value called OC. The steps below are adapted from the guidance on the National Law Information Joint-Use site (open.law.go.kr) as of October 2026.

1. Sign up at open.law.go.kr. Your login ID is your email address.

2. Under 'OPEN API 신청' ('Apply for OPEN API'), choose the data you want to use and submit a usage application. Requests for a type you have not applied for fail authentication, so choose court decisions and legal interpretations along with statutes.

3. When applying, register the IP (the address that identifies that computer on the internet) of the PC or server that will call the API. If you will call it from a website, register that domain too. Calling from an unregistered IP returns an authentication failure notice, and even after approval you can add IPs from the application list.

4. Once the administrator reviews and approves it, you can start using it. According to the site's guidance, applications are processed within 1 to 2 days.

5. Your OC value is the part of your signup email ID before the @. You call the API by adding OC=value to the request URL. Example: https://www.law.go.kr/DRF/lawSearch.do?OC=<your-id>&target=prec&type=JSON&search=2&query=<keyword>

Do not write the OC value into prompts or code; keep it in an environment variable such as LAW_API_OC. An environment variable is a configuration value, kept outside the code, that a program reads when it runs.