The Gauntlet Loop with Claude Code
Automate a builder/critic agent loop that iterates code until it passes a quality rubric.
Introduction
30 Second Summary
AI can generate a working web page in seconds. But "it works" and "it's good" are two very different standards.
In this project, you will define a quality bar and build a bash script that orchestrates two Claude Code agents in a loop. The Builder agent generates code, the Critic agent grades it, and the loop keeps running until every criterion passes.
What You'll Build
Run a single command and watch your terminal cycle through rounds of critique and repair until a rough landing page transforms into one that passes every quality standard you set.
By the end of this project, you'll have:
- A quality rubric with six demanding criteria that define what "good" looks like for any AI-generated page.
- Two Claude Code agents that take turns in your terminal. A Builder generates and fixes code while a read-only Critic grades it against your rubric.
- A gauntlet script that wires Builder and Critic into an automated feedback loop, iterating until every criterion passes or a safety cap is reached.
- Secret Mission: A Lighthouse performance gate that blocks the loop until the page scores 90+ on real browser metrics.
Are there any prerequisites?
You need Claude Code installed and authenticated with a Pro, Max, or Team plan. You also need Node.js installed on your machine.
Before We Start
Before you dive in, take a moment to lock in what you're about to build and why it matters. You are about to create an automated quality loop where two Claude Code agents push each other until the output meets a bar you define.
Set Up the Arena
Before running any agents, you need a clean workspace. A dedicated project folder with version control tracks every iteration of your landing page, making it easy to compare and roll back.
You also need a project context file. Claude Code reads a file called CLAUDE.md at the start of every session, including headless -p calls. This grounds both agents automatically without you repeating context in every prompt.
In this step, get ready to:
- Create the project folder and initialize version control.
- Write a CLAUDE.md context file for both agents.
Create the project folder and initialize Git
- Press Cmd+Space (macOS) or the Windows key (Windows) to open your search bar.
- Type Terminal and press Enter to open it.
- Move to your Desktop by running this command:
cd ~/Desktop
- Create the project folder, move into it, and initialize a Git repository by running these commands:
mkdir gauntlet-loop
cd gauntlet-loop
git init
You should see Initialized empty Git repository followed by the path to your new .git folder.
Don't see the initialized message?
- Make sure you ran cd ~/Desktop first. Without it, the folder lands wherever your terminal opened.
- Check that you typed gauntlet-loop exactly, with a hyphen between the two words.
- Help me troubleshoot my git init command.
Create CLAUDE.md
This file tells Claude Code what the project contains and what each agent is allowed to do. Because Claude Code loads CLAUDE.md automatically, every headless -p call starts with the same shared context.
- Create a new file called CLAUDE.md inside the gauntlet-loop folder with the following content:
# Gauntlet Loop Project
Build and refine a landing page through an automated builder/critic loop.
## Structure
- quality-bar.md: The rubric every build must pass
- index.html: The landing page (generated and refined by the Builder agent)
- gauntlet.sh: The orchestration script that runs the builder/critic loop
## Roles
- Builder: generates and edits index.html (uses Read, Edit, Bash tools)
- Critic: grades index.html against quality-bar.md (uses Read tool only)
What does CLAUDE.md do?
The ## Structure section maps the three files both agents will work with. This prevents the Builder from creating unexpected files and keeps the Critic focused on the right targets.
The ## Roles section defines what each agent can do. The Builder gets Read, Edit, and Bash tools. The Critic gets Read only.
- Save CLAUDE.md.
- Confirm the file exists by running:
ls
You should see CLAUDE.md listed in the output.
Don't see CLAUDE.md?
- Make sure you saved the file inside the gauntlet-loop folder, not in a parent directory.
- Check that the filename is exactly CLAUDE.md with all caps and the .md extension.
- Help me figure out why my CLAUDE.md file is not showing up.
✔️ Awesome, I've got everything!
Great. Double check you've saved CLAUDE.md before moving on.
ⓧ I'd like to double check the full code
Here is the complete CLAUDE.md file. Compare it to yours and make sure every line matches.
# Gauntlet Loop Project
Build and refine a landing page through an automated builder/critic loop.
## Structure
- quality-bar.md: The rubric every build must pass
- index.html: The landing page (generated and refined by the Builder agent)
- gauntlet.sh: The orchestration script that runs the builder/critic loop
## Roles
- Builder: generates and edits index.html (uses Read, Edit, Bash tools)
- Critic: grades index.html against quality-bar.md (uses Read tool only)
Verify your setup
Before you run the next command, what do you expect to see printed to the terminal?
- Confirm Claude Code is accessible from inside the project by running:
claude --version
You should see a version number printed to the terminal. This confirms Claude Code is installed and accessible from inside the gauntlet-loop directory.
Don't see a version number?
- Claude Code may not be on your PATH. Try closing and reopening your terminal, then running the command again.
- If the command is still not found, reinstall Claude Code by running:
curl -fsSL https://claude.ai/install.sh | bash
- Help me get Claude Code working in my terminal.
Your arena is set. Next, you'll define the quality bar that every build must survive and generate a deliberately rough first landing page.
Define the Quality Bar and Build v1
Your arena is ready. The gauntlet-loop folder has version control and a CLAUDE.md that will ground both agents in every headless call.
Most projects start with code. This one starts with a question: what does "good" look like? You will write a quality rubric before generating a single line of code, then run the Builder with a deliberately vague prompt to see what happens when code is built without aiming at the bar.
In this step, get ready to:
- Write a quality-bar.md rubric with 6 demanding criteria.
- Run the Builder agent with a deliberately vague prompt to generate index.html.
- Open the page in your browser and see the result.
Write the quality bar rubric
The rubric is the foundation of the entire gauntlet. It defines what "pass" looks like across 6 criteria, so the Critic agent has a concrete bar to enforce and the Builder has a clear target to aim for.
Each criterion is designed to evaluate intent and quality, not just the presence of an HTML tag. A quick, prompt-minimal build will reliably miss 3 to 5 of them.
- Create a file called quality-bar.md in the gauntlet-loop folder and add the rubric header and all 6 criteria:
# Quality Bar for Beacon Roasters Landing Page
You are a strict code reviewer. Read index.html and grade it against every criterion below. For each criterion, describe what you found in the code and compare it to what the criterion requires. Partial compliance is a fail.
## Criteria
1. Structural semantics: The page uses a deliberate landmark hierarchy. There is exactly one main element. The header, main, and footer elements divide the page into three top-level regions. Within main, each content block is wrapped in a section element with a heading that describes it. Generic div elements are not used where a semantic element exists.
2. Responsive craft: The page includes a viewport meta tag AND at least one CSS media query that creates a visible layout change at a specific breakpoint (e.g., switching from a single column to a multi-column grid). Simply having a viewport meta tag without a media query that alters layout is a fail.
3. Accessibility depth: The page includes a skip-navigation link as the first focusable element that jumps to main content. All interactive elements (links, buttons) have a visible focus outline using :focus-visible. Every image has a descriptive, non-empty alt attribute. If a form exists, every input has a programmatically associated label element.
4. Self-contained assets: The page makes zero external network requests. No link, script, or img element references an external URL (http:// or https://). All CSS is in a style element or inline. All images use inline SVG, data URIs, or CSS-generated shapes. External Google Fonts, CDN-hosted frameworks, and remotely loaded icons are all fails.
5. Content substance: The page includes all of the following content sections: (a) a hero with a heading, a supporting sentence, and a call-to-action button; (b) a features section displaying at least 3 distinct items in a grid or card layout, each with a heading and description; (c) a testimonial section with a quoted review, a customer name, and a detail that makes the attribution feel real (e.g., "regular since 2019"); (d) a footer containing a copyright notice with the current year.
6. Interaction and motion: At least one interactive element (button or link) has a CSS transition property that animates a visual change on hover (e.g., background color, transform, box-shadow). The page also includes a prefers-reduced-motion media query that disables or reduces the transition for users who have requested reduced motion.
Why are these criteria so specific?
Each criterion asks the Critic to evaluate quality, not just check for a tag. For example, criterion 2 does not pass just because a viewport meta tag exists. It requires a CSS media query that creates a visible layout shift at a breakpoint.
This specificity is what makes the rubric useful. A vague prompt like "include some basic content and styling" almost never produces skip-navigation links, ARIA landmarks, named testimonials, or prefers-reduced-motion media queries. The criteria are calibrated to catch that gap.
- Now add the verdict format section below the criteria in quality-bar.md:
## Verdict Format
After evaluating all 6 criteria, output your verdict in exactly one of these two formats:
If all 6 criteria pass:
VERDICT: PASS
If any criterion fails:
VERDICT: FAIL
For each failing criterion, output a numbered item with the criterion name, what you found in the code, and what was missing or wrong. Do not list passing criteria in the failure list.
Why a machine-parseable verdict format?
The exact strings VERDICT: PASS and VERDICT: FAIL are not just for humans. In Step 4, you will write a bash script that uses grep to detect the verdict and decide whether to loop or stop. A predictable format makes that automation possible without JSON parsing.
- Save quality-bar.md.
- Confirm the file exists by running:
ls
You should see CLAUDE.md and quality-bar.md listed in the output.
✔️ Awesome, I've got everything!
Great. Make sure you have saved quality-bar.md before moving on.
ⓧ I'd like to double check the full code
Here is the complete quality-bar.md file. Compare it against yours and make sure all 6 criteria and the verdict format section are present.
# Quality Bar for Beacon Roasters Landing Page
You are a strict code reviewer. Read index.html and grade it against every criterion below. For each criterion, describe what you found in the code and compare it to what the criterion requires. Partial compliance is a fail.
## Criteria
1. Structural semantics: The page uses a deliberate landmark hierarchy. There is exactly one main element. The header, main, and footer elements divide the page into three top-level regions. Within main, each content block is wrapped in a section element with a heading that describes it. Generic div elements are not used where a semantic element exists.
2. Responsive craft: The page includes a viewport meta tag AND at least one CSS media query that creates a visible layout change at a specific breakpoint (e.g., switching from a single column to a multi-column grid). Simply having a viewport meta tag without a media query that alters layout is a fail.
3. Accessibility depth: The page includes a skip-navigation link as the first focusable element that jumps to main content. All interactive elements (links, buttons) have a visible focus outline using :focus-visible. Every image has a descriptive, non-empty alt attribute. If a form exists, every input has a programmatically associated label element.
4. Self-contained assets: The page makes zero external network requests. No link, script, or img element references an external URL (http:// or https://). All CSS is in a style element or inline. All images use inline SVG, data URIs, or CSS-generated shapes. External Google Fonts, CDN-hosted frameworks, and remotely loaded icons are all fails.
5. Content substance: The page includes all of the following content sections: (a) a hero with a heading, a supporting sentence, and a call-to-action button; (b) a features section displaying at least 3 distinct items in a grid or card layout, each with a heading and description; (c) a testimonial section with a quoted review, a customer name, and a detail that makes the attribution feel real (e.g., "regular since 2019"); (d) a footer containing a copyright notice with the current year.
6. Interaction and motion: At least one interactive element (button or link) has a CSS transition property that animates a visual change on hover (e.g., background color, transform, box-shadow). The page also includes a prefers-reduced-motion media query that disables or reduces the transition for users who have requested reduced motion.
## Verdict Format
After evaluating all 6 criteria, output your verdict in exactly one of these two formats:
If all 6 criteria pass:
VERDICT: PASS
If any criterion fails:
VERDICT: FAIL
For each failing criterion, output a numbered item with the criterion name, what you found in the code, and what was missing or wrong. Do not list passing criteria in the failure list.
Run the Builder with a deliberately vague prompt
Now you will generate the first version of the landing page. The prompt you give the Builder is deliberately vague. It says nothing about accessibility, responsiveness, testimonials, transitions, or any of the 6 rubric criteria.
- Generate index.html by running Claude Code in headless mode with this command:
claude -p "Create a single-file landing page for a coffee shop called Beacon Roasters in index.html. Include some basic content and styling." \
--allowedTools "Read,Edit,Bash" \
--output-format text
What does this command do?
- -p runs Claude Code in headless (print) mode. It executes the prompt and exits without opening an interactive session.
- --allowedTools "Read,Edit,Bash" gives the Builder permission to read files, edit files, and run bash commands. In headless mode, tools not in this list are denied automatically since there is no interactive prompt to approve them.
- --output-format text prints plain text output to the terminal instead of JSON.
The command takes 15 to 45 seconds to run. You will see Claude Code's progress in the terminal as it creates and writes the file.
- Confirm the file was created by running:
ls
You should now see CLAUDE.md, index.html, and quality-bar.md in the output.
Don't see index.html?
- Make sure you are signed into Claude Code. Run claude in interactive mode to check your authentication status, then exit and retry the headless command.
- Check that your terminal is in the gauntlet-loop directory. Run pwd to confirm.
- If the command timed out or errored, run it again. Headless calls occasionally fail on first attempt due to model load.
Why can't Claude Code create my index.html file?
Open the page in your browser
Before you look at the page, take a moment to think. You gave the Builder a prompt that said "some basic content and styling" and nothing else. What do you expect to see?
- Open the landing page in your browser by running:
open index.html
You should see a working landing page for Beacon Roasters with styled content. It has a heading, some text, and a reasonable visual layout.
It looks fine. Is that the point?
The page works. It looks reasonable. But "works" and "meets the bar" are not the same thing.
You wrote 6 specific criteria covering skip-navigation links, responsive breakpoints, named testimonials, CSS transitions, and reduced-motion queries. The vague Builder prompt mentioned none of them. That gap between what looks good in the browser and what the rubric actually requires is what the rest of this project exists to close.
The page looks finished. But the rubric says otherwise. In the next step, you will unleash the Critic agent and see exactly where this page falls short.
Unleash the Critic
Your landing page is live in the browser, and it looks perfectly reasonable. But "looks reasonable" and "meets the bar" are two different things.
You wrote 6 specific criteria in quality-bar.md that define what "good" actually means. In this step, you'll run a Critic agent that can only read files. It cannot edit anything. Its sole job is to judge.
In this step, get ready to:
- Run the Critic agent with read-only tool access against the rubric.
- Read the failure report and identify which criteria the page misses.
Run the Critic agent
The Critic is a second Claude Code agent with a crucial constraint. Its --allowedTools flag is set to "Read" only. It can read quality-bar.md and index.html, but it cannot edit, create, or delete anything.
Before you run this, what do you think the Critic will find? The page looks fine visually. Will it pass all 6 criteria, or will some fall short?
- Run the Critic agent against your rubric and landing page by running this command in your terminal:
claude -p "Read quality-bar.md for the grading rubric. Then read index.html and grade it against every criterion. For each criterion, explain what you found in the code and whether it meets the bar. Be strict: partial compliance is a fail. Output your verdict in the exact format specified in quality-bar.md." \
--allowedTools "Read" \
--output-format text
What do these flags do?
- -p runs Claude Code in headless print mode. It processes the prompt and outputs the response directly to your terminal, with no interactive session.
- --allowedTools "Read" restricts the agent to the Read tool only. In -p mode, any tool not in this list is denied since there is no interactive prompt to approve it.
- --output-format text returns plain text output instead of JSON or streaming JSON.
This separation of concerns is the core of the pattern. The agent that judges is not the agent that builds.
The Critic outputs VERDICT: FAIL followed by a numbered list of every criterion the page misses. The page that looked fine in the browser just failed the quality bar.
Not seeing any output?
- Make sure you are in the gauntlet-loop directory. The Critic needs to find both quality-bar.md and index.html in the current directory.
- Check that Claude Code is authenticated by running claude --version first. If it prompts you to sign in, complete the sign-in before retrying.
- The command may take 15-45 seconds to complete. Wait for the full output before assuming it failed.
Still stuck? Help me debug why my Critic command is not producing output.
Read the failure report
Scroll through the Critic's output. Because the Builder prompt in Step 2 never mentioned any of the 6 criteria, the page has real gaps. Common failures include:
- No skip-navigation link or :focus-visible outlines on interactive elements.
- No CSS media query that creates a visible layout shift at a breakpoint.
- No named testimonial with a real attribution detail.
- No CSS transition on interactive elements, and no prefers-reduced-motion media query.
- External font or CDN references that violate the self-contained assets criterion.
Why is this feedback different from a checklist?
The Critic does not just list missing tags. For each failing criterion, it explains what it found in the code, compares it to what the rubric requires, and names the specific gap.
This makes the feedback actionable. Instead of "add a skip link," the Critic says what the page has, what the rubric demands, and why the current code falls short. This is the feedback that will drive the Builder's fixes in the next step.
Your exact list of failures will vary depending on what the Builder generated. The key observation is that 3 or more criteria fail, proving the gap between a page that "works" and a page that "meets the bar."
The gap is visible. The page works, but it does not meet the bar you defined. Next up, you'll wire the Critic and Builder into an automated loop so the Builder fixes every failing criterion without you copying feedback by hand.
Wire and Run the Gauntlet
The Critic did its job. You now have a detailed failure report listing exactly where index.html falls short of the quality bar you defined.
Copying that feedback into a Builder prompt by hand works once, but it does not scale. In this step, you will write a bash script that wires the Critic and Builder into an automated loop. The Critic's output becomes the Builder's input, and the cycle repeats until the page passes every criterion or a safety cap is reached.
In this step, get ready to:
- Write gauntlet.sh to orchestrate the Critic/Builder feedback loop.
- Run the gauntlet and watch the loop iterate until the page passes.
- Refresh the browser to see the improved landing page.
Write the orchestration script
The gauntlet script has one job: run the Critic, check the verdict, and if it fails, feed the feedback straight into the Builder. Then repeat.
- Create a new file called gauntlet.sh in your gauntlet-loop folder and add the first section of the script:
#!/bin/bash
set -euo pipefail
MAX_LOOPS=5
PASS_MARKER="VERDICT: PASS"
echo "=== THE GAUNTLET ==="
echo "Max iterations: $MAX_LOOPS"
echo ""
for i in $(seq 1 $MAX_LOOPS); do
echo "--- Iteration $i/$MAX_LOOPS: Critic ---"
FEEDBACK=$(claude -p "Read quality-bar.md for the grading rubric. Then read index.html and grade it against every criterion. For each criterion, explain what you found in the code and whether it meets the bar. Be strict: partial compliance is a fail. Output your verdict in the exact format specified in quality-bar.md." \
--allowedTools "Read" \
--output-format text)
echo "$FEEDBACK"
echo ""
if echo "$FEEDBACK" | grep -q "$PASS_MARKER"; then
echo "=== GAUNTLET PASSED on iteration $i ==="
exit 0
fi
What does this code do?
- set -euo pipefail enables bash strict mode. The script exits immediately on any error, treats unset variables as errors, and catches failures inside piped commands.
- MAX_LOOPS=5 sets a safety cap so the loop cannot run forever if the page never passes.
- The for loop runs the Critic agent using the exact same read-only claude -p command from the previous step, captures its full output into the FEEDBACK variable, and prints it.
- grep -q checks the Critic's output for the VERDICT: PASS string. If found, the script prints a success message and exits.
- Still in gauntlet.sh, add the following code directly below the fi line:
echo "--- Iteration $i/$MAX_LOOPS: Builder ---"
claude -p "You are a web developer. Read quality-bar.md for the full quality bar. The Critic found the following problems with index.html. Fix every failing criterion. Edit index.html directly. Do not remove content that already passes.
CRITIC FEEDBACK:
$FEEDBACK" \
--allowedTools "Read,Edit,Bash" \
--output-format text
echo "Builder completed fixes."
echo ""
done
echo "=== GAUNTLET FAILED after $MAX_LOOPS iterations ==="
exit 1
What does this section do?
When the Critic returns VERDICT: FAIL, the script falls through to the Builder. The entire Critic feedback is injected into the Builder's prompt via the $FEEDBACK variable, so the Builder knows exactly what to fix.
The Builder uses --allowedTools "Read,Edit,Bash" so it can read files, edit index.html, and run commands. The Critic only had --allowedTools "Read". This separation is enforced at the tool level, not just the prompt level.
After the Builder finishes, the loop starts a new iteration and the Critic grades the updated page. If the safety cap of 5 is reached without a pass, the script exits with a failure message.
- Save gauntlet.sh.
- Confirm the file exists by running:
ls gauntlet.sh
You should see gauntlet.sh listed in the output.
Don't see gauntlet.sh?
- Make sure you saved the file in the gauntlet-loop directory, not a parent or nested folder.
- Check that your terminal is still in the gauntlet-loop directory by running pwd.
- help me find where my gauntlet.sh file was saved
✔️ Awesome, I've got everything!
Great. Double check you have saved gauntlet.sh before moving on.
ⓧ I'd like to double check the full code
The full gauntlet.sh file is built from two sections. Scroll up and confirm each one is present in order:
- Section 1: starts with #!/bin/bash and ends with fi. This covers the shebang, variables, the for loop, the Critic call, and the pass check.
- Section 2: starts with the Builder echo line and ends with exit 1. This covers the Builder call, the done keyword closing the loop, and the failure exit.
Make sure the two sections are directly adjacent with no gap or duplication between the fi at the end of section 1 and the echo at the start of section 2.
Run the gauntlet
The script is ready. Before running it, you need to make it executable.
- Make the script executable and run it:
chmod +x gauntlet.sh
bash gauntlet.sh
What does chmod +x do?
chmod +x adds execute permission to the file. Without it, your shell would refuse to run the script directly. Using bash gauntlet.sh explicitly invokes bash, but the permission is good practice for later runs.
Before you watch the output scroll by, take a moment to predict: do you think the page will pass on the first Critic iteration, or will it take a few rounds? Why?
Each iteration prints the Critic's verdict followed by the Builder's fix summary. The loop typically passes in 2-3 iterations. You will see something like this at the end:
=== GAUNTLET PASSED on iteration 3 ===
What just happened?
Each claude -p call runs Claude Code in headless mode. The Critic read quality-bar.md and index.html, produced a failure report, and the script fed that report directly into the Builder's prompt. The Builder edited index.html to fix the failing criteria, and then the Critic graded the updated page again.
This closed feedback loop is the core of the builder/critic pattern. The agent that judges never edits. The agent that edits always receives specific, structured feedback about what to fix.
🙋♀️ Gauntlet failed after 5 iterations?
If the script exits with GAUNTLET FAILED after 5 iterations, the Builder did not fully satisfy all 6 criteria within the safety cap. This can happen if one criterion keeps slipping while another is fixed.
- Run bash gauntlet.sh again. The Builder starts from the current (partially improved) index.html, so a second run often finishes what the first started.
- Read the last Critic output to see which criterion is still failing. If it is the same one repeatedly, the rubric wording for that criterion may be ambiguous. Check quality-bar.md for clarity.
- help me debug why my gauntlet loop is not passing after multiple runs
Refresh the browser and compare
- Switch back to your browser and reload the index.html page.
Compare this version to the rough v1 you saw in Step 2. The page should now have visible differences: a skip-navigation link, a responsive layout that shifts at a breakpoint, a named testimonial with real attribution, hover transitions on buttons, and no external font or CDN references.
Why does it look so different?
The v1 page was generated from a deliberately vague prompt with no mention of accessibility, responsiveness, or content requirements. The gauntlet forced the Builder to address every gap the Critic found. Each iteration layered in the missing semantics, transitions, and content.
The rubric is what drove the improvement, not a better prompt. The same vague Builder prompt produced a passing page once the loop had enough iterations to close every gap.
Your automated loop just pushed a rough draft through 6 demanding criteria and came out the other side. Next up, you will verify the result independently to make sure the pass is real and not an artifact of the script's control flow.
Verify the Result
Your gauntlet loop ran to completion and the terminal shows a passing verdict. But how do you know the result is real? The script's grep check, variable assignments, and conditional branches could all mask a false positive.
Running the Critic as a standalone command, completely outside the loop, strips away all that machinery and confirms the result independently. Then you'll lock everything into git so the passing state is preserved.
In this step, get ready to:
- Re-run the Critic independently to confirm all 6 rubric criteria pass.
- Commit all project files to git.
Re-run the Critic independently
Inside gauntlet.sh, the Critic's output flowed through variables, grep checks, and loop control flow before the script declared a pass. An independent run removes all of that machinery.
Before you run this, what do you expect the Critic to say? Will it still pass when it runs outside the gauntlet loop?
- Re-run the exact same Critic command from earlier by running:
claude -p "Read quality-bar.md for the grading rubric. Then read index.html and grade it against every criterion. For each criterion, explain what you found in the code and whether it meets the bar. Be strict: partial compliance is a fail. Output your verdict in the exact format specified in quality-bar.md." \
--allowedTools "Read" \
--output-format text
The Critic outputs VERDICT: PASS with all 6 criteria confirmed.
Why verify outside the loop?
This is the same principle as running tests outside of CI. If the Critic still says PASS with no script wrapping it, the result belongs to the page, not to the script's control flow.
The same rubric, the same agent, the same strictness. The only difference is that it ran standalone.
🙋♀️ Seeing VERDICT: FAIL?
If the Critic fails one or two criteria on this independent run, the gauntlet may have exited early due to a grep match on a partial line. Re-run the gauntlet loop to let it iterate again:
bash gauntlet.sh
After it passes, re-run the standalone Critic command above to confirm.
Still stuck? Help me debug why the Critic passes inside the loop but fails standalone.
Commit the result
With the independent verification done, lock the passing state into version control so you have a clean snapshot of the finished project.
- Stage and commit all project files by running:
git add -A && git commit -m "gauntlet: landing page passes all quality bar criteria"
- Confirm the commit exists by running:
git log --oneline
You should see your commit message gauntlet: landing page passes all quality bar criteria at the top of the log.
Seeing a git config error?
If git asks you to set your name and email before committing, run these two commands with your own details:
git config user.name "Your Name"
git config user.email "your@email.com"
Then re-run the git add -A && git commit command above.
Still stuck? Help me fix my git commit error.
You now have a reusable pattern: define a rubric, run a vague first build, then loop until the rubric passes. The rubric, the Builder prompt, and the Critic prompt are all plain text files you can adapt for any project.
Secret mission
Add a Lighthouse Performance Gate
Your gauntlet catches code quality issues through LLM judgment, but what about raw performance? In this extension, you will add a Lighthouse gate that measures your page against real browser metrics. The gauntlet will not pass until both the Critic AND Lighthouse agree the page is good.
Clean Up Your Resources
Clean Up Your Resources
This project runs entirely locally, so there are no ongoing cloud costs or running services to worry about. Decide whether to keep your project files, or delete them entirely.
Resources you used:
- gauntlet-loop project folder on your Desktop, containing CLAUDE.md, quality-bar.md, index.html, gauntlet.sh, and a local .git repository.
- report.json Lighthouse report file inside the project folder (generated during the Secret Mission).
- check-lighthouse.mjs score-checking script inside the project folder (created during the Secret Mission).
Keep everything running
No action needed. Choose this if you want to keep experimenting with the gauntlet loop, swap the rubric for a different project, or refine the Builder and Critic prompts.
- Your gauntlet-loop folder stays on your Desktop with all files intact.
- No processes are running. The gauntlet script starts and stops the HTTP server within each execution, so nothing is left behind.
- Claude Code and Node.js remain installed on your machine from before this project.
Pause - I'll come back to this later
There are no running processes or cloud services to pause. Your project files sit on disk and use minimal storage.
- The gauntlet-loop folder is ready to pick up whenever you return. Run bash gauntlet.sh from inside the folder to re-run the loop at any time.
- If you want to free up a small amount of disk space, delete report.json. The gauntlet regenerates it on each run.
Delete - I don't want to use this again
Remove all project files from your machine. This deletes everything you built in this project.
- Delete the entire gauntlet-loop folder by running this command from your Desktop:
macOS
cd ~/Desktop
rm -rf gauntlet-loop
Windows
cd $HOME\Desktop
Remove-Item -Recurse -Force gauntlet-loop
- Confirm the folder is gone by running:
ls ~/Desktop | grep gauntlet
You should see no output. The project folder and all its contents (including the git history, rubric, landing page, gauntlet script, and Lighthouse report) are gone.
Nice Work!
Nice Work!
Nice work! You just built an automated quality gauntlet that takes AI-generated code from "it works" to "it's good" without any manual copy-pasting between agents.
You've learned how to:
- Define a quality rubric (quality-bar.md) with 6 demanding criteria that evaluate intent and craft, then use it to drive an automated feedback loop.
- Orchestrate two distinct Claude Code agents in headless mode using -p and --allowedTools to enforce a strict Builder/Critic separation of concerns at the tool level.
- Write a gauntlet.sh bash script that wires the Critic/Builder feedback loop together with a safety cap, passing failure feedback directly into the next Builder iteration until the rubric passes.
- Secret Mission: Added a Lighthouse performance gate that combines subjective LLM judgment with objective machine measurement, so the gauntlet requires both to pass before the code ships.
Ready to quiz yourself?