Build an Adaptive Code Review Router
Build a policy-driven AI router for adaptive Claude Code reviews.
Introduction
30 Second Summary
Some code changes look large but carry little risk. Others change only three lines yet decide who can access sensitive information.
In this project, you will build a terminal pipeline that routes each code change to a matching Claude Code reviewer. Jev supplies typed evidence while deterministic Python policy owns the final decision.
What You'll Build
When you paste a unified diff into the finished router, the terminal explains its chosen review depth before showing the matching review.
By the end of this project, you'll have:
- An adaptive review router that sends routine changes through a concise review while reserving deeper reviews for higher-stakes code.
- A deterministic safety policy that forces credential, payment, or destructive database changes into the high-risk path.
- A machine-readable evaluation report that reveals routing accuracy, escalation mistakes, reviewer usage, resolved model usage, and estimated cost.
- Secret Mission: Make a three-line authorization diff fail closed into the high-risk reviewer despite deliberately misleading low-risk Jev signals.
Are there any prerequisites?
You should be comfortable reviewing Python code from a terminal. You also need Claude Code access on a Mac.
Before We Start
Before the hands-on work begins, lock in the rule that guides this router. AI supplies typed evidence about complexity, consequence, and confidence. Deterministic Python owns every irreversible routing decision.
Set Up the Local Project
Your policy invariant needs a dependable local foundation before any live routing call can run.
An isolated Python environment pins the TypeSafe AI Python SDK to version 0.7.2. A current Claude Code login makes the later reviewer dispatch available.
A session-only TypeSafe key gives Jev access without writing the credential into your project files.
In this step, get ready to:
- Create a local project with an active Python 3.10 or newer virtual environment.
- Install TypeSafe AI Python SDK version 0.7.2 from requirements.txt.
- Prepare authenticated access for TypeSafe and Claude Code.
Create the isolated project environment
A virtual environment keeps this project's packages separate from other Python projects on your Mac. The version check comes first because the SDK requires Python 3.10 or newer.
- Press Cmd+Space to open Spotlight.
- Type Terminal into Spotlight.
- Press Enter to open Terminal.
- Move to your Desktop by running this command:
cd ~/Desktop
What does this command do?
The cd command moves your shell to the Desktop. This gives the new project a predictable location.
- Create the jev-review-router folder by running these commands:
mkdir -p jev-review-router
cd jev-review-router
pwd
How does the folder setup work?
- The mkdir command creates the project folder on your Desktop.
- The cd command moves the shell into that folder.
- The final command prints the current location so you can confirm where later files are created.
You should see a path ending in Desktop/jev-review-router.
- Check whether your installed Python meets the project requirement by running:
python3 -c 'import sys; assert sys.version_info >= (3, 10), sys.version'
What does the Python check prove?
The assertion reads Python's version information. Returning to the prompt confirms that the interpreter is version 3.10 or newer.
✔️ The Python check passes
Your Python interpreter meets the minimum version. Continue with the environment creation below.
ⓧ The version assertion fails
Your current interpreter is older than Python 3.10. Install a newer release before creating the environment.
- Visit the official macOS Python downloads page.
- Install Python 3.10 or newer using the macOS installer.
- Return to the Terminal window from earlier.
- Repeat the version check by running:
python3 -c 'import sys; assert sys.version_info >= (3, 10), sys.version'
What confirms the upgrade?
The check returns to the prompt when the new interpreter meets the requirement. You can then continue with the project environment.
Still seeing the older interpreter?
Close the current Terminal window after installation. Open a fresh Terminal window so the shell can discover the new Python installation.
If the old interpreter still runs, ask for help identifying the newer Python executable on your Mac.
ⓧ Command not found
Your shell cannot find Python. Install Python before creating the project environment.
- Visit the official macOS Python downloads page.
- Install Python 3.10 or newer using the macOS installer.
- Return to the Terminal window from earlier.
- Repeat the version check by running:
python3 -c 'import sys; assert sys.version_info >= (3, 10), sys.version'
What confirms the installation?
The check returns to the prompt when the shell can find a compatible interpreter. That interpreter is ready to create the isolated environment.
Python still unavailable?
Open a fresh Terminal window after the installer finishes. A new shell session can detect the updated command path.
If the command is still unavailable, ask for help checking your Python installation path.
- Create the local .venv environment by running:
python3 -m venv .venv
What does this command create?
Python creates an isolated interpreter inside the .venv folder. Packages installed after activation stay scoped to this project.
- Activate the environment in the current terminal by running:
source .venv/bin/activate
What does activation change?
Activation makes the environment's Python interpreter the default for this shell. Later installation commands now target .venv.
You should see (.venv) near the start of your terminal prompt.
- Install the pinned TypeSafe dependency into the active environment by running:
python -m pip install typesafe-sdk==0.7.2
Why pin the SDK version?
The version pin keeps the SDK interface consistent with the router code you build later. Every learner works against the same dependency surface.
The installation output should finish successfully with TypeSafe SDK version 0.7.2 available inside .venv.
Package installation failing?
Confirm that (.venv) appears in your prompt. A missing prefix means the environment is inactive.
If the installation still fails, ask for help diagnosing the TypeSafe SDK installation.
The project now needs a dependency file that another developer can inspect. You will create it in Visual Studio Code with the exact SDK pin.
- Open the current jev-review-router folder in Visual Studio Code by running:
code .
What does the editor command do?
The code command opens the current folder as a Visual Studio Code workspace. Files created in the editor now land inside jev-review-router.
You should see the jev-review-router folder in the left sidebar.
Editor command unavailable?
Use Visual Studio Code's official macOS setup guide to add the editor command to your shell path.
If the command remains unavailable, ask for help connecting Visual Studio Code to your terminal.
- Select the new-file icon beside the jev-review-router folder in the left sidebar.
- Name the new file requirements.txt.
- Paste this dependency pin into requirements.txt:
typesafe-sdk==0.7.2
What does this file record?
The file declares the exact Python package required by the router. Its version matches the package already installed in .venv.
- Press Cmd+S to save requirements.txt.
- Confirm that the left sidebar lists requirements.txt beside the .venv folder.
Dependency file missing?
Confirm that the editor title shows the jev-review-router workspace. Creating the file in another workspace places it outside this project.
If the file is still missing, ask for help creating requirements.txt in the correct folder.
✔️ Awesome, I've got everything!
Great. Double-check that requirements.txt is saved before you continue.
ⓧ I'd like to double check the full code
typesafe-sdk==0.7.2
Create and export the TypeSafe API key
Credentials need a narrow boundary because they authorize metered requests. You will keep this key in the current shell instead of saving it in requirements.txt or another project file.
Jev requests are metered
TypeSafe does not document public free credits for these calls. Jev 1.13 costs $0.042 per million input tokens.
Output tokens are free. The calls in this project are small.
- Visit the TypeSafe API keys page.
- Sign in with your TypeSafe account if prompted.
- Create a TypeSafe API key from the keys page.
- Copy the new key to your clipboard.
- Return to the Terminal window from earlier.
- Export the key in the current terminal by replacing your-key-here before you run this command:
export TYPESAFE_API_KEY="your-key-here"
Why export the key this way?
The TYPESAFE_API_KEY environment variable becomes available to commands launched from this shell. Closing the Terminal window removes the value from that session.
The SDK reads this variable automatically when the router creates its TypeSafe client.
That sensitive step is safely contained. The key stays out of your project files.
Key export not working?
Keep the double quotes around the key. Replace only the placeholder text inside them.
If later TypeSafe requests reject the credential, ask for help checking the shell-scoped export without sharing the key.
Update and verify Claude Code
The router later invokes Claude Code through its local command-line interface. Updating the installed client first ensures that the project can use current non-interactive review features.
- Update Claude Code from the active terminal by running:
claude update
What does the update command do?
Claude Code checks for the latest available client release. Keeping the client current prepares it for the project-scoped reviewer profiles you add later.
Claude Code update failing?
Confirm that the terminal can access the internet. An interrupted download can prevent the update from completing.
If the update still fails, ask for help diagnosing your Claude Code installation.
- Check the stored Claude Code authentication state by running:
claude auth status
What does the authentication check prove?
Claude Code prints its authentication status as structured output. A successful authenticated result confirms that the later reviewer subprocesses can use your existing access.
✔️ Claude Code is authenticated
Your existing Claude Code login is ready. Continue to the final environment check.
ⓧ Claude Code is not authenticated
Claude Code needs an authenticated session before the router can dispatch reviews.
- Start the Claude Code sign-in flow by running:
claude auth login
What happens during sign-in?
Claude Code starts its authentication flow for your existing access. Complete the prompts before returning to the terminal.
- Confirm the new authentication state by running:
claude auth status
What confirms the login?
A successful authenticated result confirms that Claude Code can accept later review requests from the router.
Authentication still incomplete?
Complete every browser prompt opened by the sign-in flow. Return to the same terminal after access is approved.
If authentication still fails, ask for help connecting your Claude Code access.
Before you run the final check, which three signals will prove that the interpreter, SDK, and Claude Code session are ready?
- Verify the complete local setup by running these checks:
python --version
python -c "import typesafe_sdk; print('TypeSafe SDK ready')"
claude auth status
What does the final check cover?
- The first check reports the interpreter used by the active virtual environment.
- The second check imports the installed SDK before printing its readiness message.
- The third check confirms that Claude Code has an authenticated session.
You should see Python 3.10 or newer. You should also see TypeSafe SDK ready.
Claude Code should report an authenticated session. Together these results prove that the local foundation is ready.
One of the final checks failing?
Confirm that (.venv) still appears in the terminal prompt. Reactivate the environment if the Python version or SDK import is wrong.
If the Claude Code check fails, complete the authentication path above. Ask for help interpreting any remaining setup failure.
Your local review workspace is ready to make metered Jev decisions and launch authenticated Claude Code reviews. Next up, you will expose why one fixed review depth fails across changes with very different consequences.
Expose the Single-Reviewer Problem
Your isolated environment is ready. The project can now run a controlled routing experiment.
A single review depth treats every change alike. This step measures where that shortcut wastes review effort or misses required depth.
In this step, get ready to:
- Create seven labeled unified diffs with expected review routes.
- Build a baseline that predicts complex for every case.
- Measure both escalation error types by running the baseline.
Create the labeled evaluation set
A labeled evaluation set gives each unified diff an expected route. Storing the cases as JSON gives the baseline a consistent input format.
- Select the file-creation icon in the VS Code file sidebar from earlier.
You will see a field where you can name the new file.
- Enter eval_cases.json in the filename field.
- Press Enter.
The empty eval_cases.json file is now open in the editor.
- Fill eval_cases.json with the seven labeled diffs by pasting this content:
[
{"id":"docs-wording","expected_route":"routine","diff":"diff --git a/README.md b/README.md\n-Run the server.\n+Run the local development server.\n"},
{"id":"local-rename","expected_route":"routine","diff":"diff --git a/items.py b/items.py\n-result = [item for item in items if item.active]\n-return result\n+active_items = [item for item in items if item.active]\n+return active_items\n"},
{"id":"retry-loop","expected_route":"complex","diff":"diff --git a/client.py b/client.py\n+for attempt in range(3):\n+ try:\n+ return send_request()\n+ except TimeoutError:\n+ if attempt == 2:\n+ raise\n"},
{"id":"schema-extension","expected_route":"complex","diff":"diff --git a/migrations/014_orders.sql b/migrations/014_orders.sql\n+ALTER TABLE orders ADD COLUMN external_id TEXT;\n+CREATE INDEX orders_external_id_idx ON orders(external_id);\n"},
{"id":"payment-boundary","expected_route":"high_risk","diff":"diff --git a/payments.py b/payments.py\n-if amount > 0:\n+if amount >= 0:\n charge_card(amount)\n"},
{"id":"destructive-migration","expected_route":"high_risk","diff":"diff --git a/migrations/015_cleanup.sql b/migrations/015_cleanup.sql\n+DROP TABLE archived_orders;\n"},
{"id":"three-line-authorization","expected_route":"high_risk","diff":"diff --git a/authorization.py b/authorization.py\n-if user.is_admin:\n+if user:\n return view_audit_log()\n"}
]
What Do the Labels Capture?
- Each id gives the case a short name for the results table.
- Each expected_route records the required review depth.
- The routine labels cover a documentation edit plus a local rename.
- The complex labels cover retry behavior plus a schema extension.
- The high_risk labels cover payments.
- The remaining high-risk labels cover destructive data operations plus authorization.
- Save eval_cases.json.
- Confirm the editor shows eval_cases.json without an unsaved-change marker.
Cases File Not Saving?
- Check that the filename is exactly eval_cases.json.
- Check that every diff remains inside its quoted JSON value.
- Compare the file with the complete reference below.
Help me troubleshoot my evaluation cases file.
✔️ Awesome, I've got everything!
Your seven labeled cases are saved in eval_cases.json.
ⓧ I'd like to double check the full code
[
{"id":"docs-wording","expected_route":"routine","diff":"diff --git a/README.md b/README.md\n-Run the server.\n+Run the local development server.\n"},
{"id":"local-rename","expected_route":"routine","diff":"diff --git a/items.py b/items.py\n-result = [item for item in items if item.active]\n-return result\n+active_items = [item for item in items if item.active]\n+return active_items\n"},
{"id":"retry-loop","expected_route":"complex","diff":"diff --git a/client.py b/client.py\n+for attempt in range(3):\n+ try:\n+ return send_request()\n+ except TimeoutError:\n+ if attempt == 2:\n+ raise\n"},
{"id":"schema-extension","expected_route":"complex","diff":"diff --git a/migrations/014_orders.sql b/migrations/014_orders.sql\n+ALTER TABLE orders ADD COLUMN external_id TEXT;\n+CREATE INDEX orders_external_id_idx ON orders(external_id);\n"},
{"id":"payment-boundary","expected_route":"high_risk","diff":"diff --git a/payments.py b/payments.py\n-if amount > 0:\n+if amount >= 0:\n charge_card(amount)\n"},
{"id":"destructive-migration","expected_route":"high_risk","diff":"diff --git a/migrations/015_cleanup.sql b/migrations/015_cleanup.sql\n+DROP TABLE archived_orders;\n"},
{"id":"three-line-authorization","expected_route":"high_risk","diff":"diff --git a/authorization.py b/authorization.py\n-if user.is_admin:\n+if user:\n return view_audit_log()\n"}
]
How to Compare This File
This reference should match your saved file character for character. The escaped newline sequences keep each unified diff inside one JSON string.
Build the fixed baseline
A small Python baseline makes the routing shortcut measurable. Its route ranks show whether a prediction sits above or below the expected depth.
Every case receives the fixed complex prediction. The resulting counters reveal the trade-off after you run the experiment.
- Select the file-creation icon in the VS Code file sidebar from earlier.
You will see the filename field again.
- Enter baseline.py in the filename field.
- Press Enter.
The empty baseline.py file is now open beside the evaluation set.
- Build the fixed-route evaluator in baseline.py by pasting this code:
from __future__ import annotations
import json
from pathlib import Path
ROUTE_RANK = {"routine": 0, "complex": 1, "high_risk": 2}
def main() -> None:
cases = json.loads(Path("eval_cases.json").read_text())
predicted = "complex"
correct = 0
over = 0
under = 0
print("case\texpected\tpredicted")
for case in cases:
expected = case["expected_route"]
correct += predicted == expected
over += ROUTE_RANK[predicted] > ROUTE_RANK[expected]
under += ROUTE_RANK[predicted] < ROUTE_RANK[expected]
print(f"{case['id']}\t{expected}\t{predicted}")
print(f"\naccuracy={correct / len(cases):.1%}")
print(f"over_escalations={over}")
print(f"under_escalations={under}")
if __name__ == "__main__":
main()
What Does the Baseline Measure?
- The json module loads the labeled evaluation cases.
- Path reads eval_cases.json from the project folder.
- ROUTE_RANK converts each route into an ordered review depth.
- correct counts exact route matches.
- over counts predictions above the expected route.
- under counts predictions below the expected route.
- Save baseline.py.
- Confirm the editor shows baseline.py without an unsaved-change marker.
Baseline File Incomplete?
- Confirm that baseline.py is inside the jev-review-router folder.
- Confirm that eval_cases.json is in the same folder.
- Check that the file ends with the main() call.
Help me troubleshoot my fixed baseline.
✔️ Awesome, I've got everything!
Your fixed evaluator is saved in baseline.py.
ⓧ I'd like to double check the full code
from __future__ import annotations
import json
from pathlib import Path
ROUTE_RANK = {"routine": 0, "complex": 1, "high_risk": 2}
def main() -> None:
cases = json.loads(Path("eval_cases.json").read_text())
predicted = "complex"
correct = 0
over = 0
under = 0
print("case\texpected\tpredicted")
for case in cases:
expected = case["expected_route"]
correct += predicted == expected
over += ROUTE_RANK[predicted] > ROUTE_RANK[expected]
under += ROUTE_RANK[predicted] < ROUTE_RANK[expected]
print(f"{case['id']}\t{expected}\t{predicted}")
print(f"\naccuracy={correct / len(cases):.1%}")
print(f"over_escalations={over}")
print(f"under_escalations={under}")
if __name__ == "__main__":
main()
How to Compare This File
This reference shows the complete baseline for this step. Matching it keeps the fixed prediction plus all three counters intact.
Run the single-reviewer experiment
The fixed rule is ready to face all seven labels. Before you run it, which labels do you think one complex prediction will over-escalate or under-escalate?
- Run the experiment in the active terminal by using this command:
python baseline.py
What Did the Baseline Reveal?
- The table contains seven case rows.
- Every predicted route is complex.
- accuracy=28.6% shows that two predictions match their labels.
- over_escalations=2 comes from the two routine changes.
- under_escalations=3 comes from the three high-risk changes.
The routine cases receive more review depth than their labels require. The consequential cases receive less depth than their labels require.
You have exposed the baseline's trade-off. One fixed reviewer tier cannot match changes with different consequences.
Baseline Table Missing?
- Switch back to the active terminal from the setup step.
- Confirm that the terminal is inside the jev-review-router folder.
- Confirm that both project files appear in the VS Code file sidebar.
- Compare each file with its complete reference above.
Help me troubleshoot the baseline output.
The single-reviewer problem is now visible in measurable results. Next, Jev supplies typed evidence for a deterministic routing policy.
Turn Jev Signals into Policy
Your single-reviewer baseline proved that one fixed tier wastes attention on routine diffs. It also leaves consequential changes without a guaranteed deep review.
Now Jev supplies typed evidence about each change. Deterministic Python policy turns that evidence into the final review route.
In this step, get ready to:
- Ask two atomic Choice questions about implementation complexity and production consequence.
- Convert the resulting signals into routine, complex, or high-risk routes using Python policy.
- Pipe a harmless documentation diff through the dry-run path to inspect the decision JSON.
Capture typed Jev evidence
A Choice question limits an answer to named options. It also preserves the probability assigned to every option.
Separate questions keep complexity distinct from consequence. A three-line payment change can be simple to understand while carrying serious production impact.
Why use two atomic questions?
One broad prompt can blur code difficulty with business impact. Two atomic questions expose both signals for inspection.
Python can then combine those signals through a stable policy. The final route remains easy to audit.
- Switch back to Visual Studio Code with the existing project folder.
- Select the new-file control in the Explorer sidebar.
- Name the new file router.py.
- Add the imports plus the project constants by pasting this code:
from __future__ import annotations
import argparse
import json
import subprocess
import sys
from dataclasses import asdict, dataclass
from pathlib import Path
from policy import HARD_GATE_TERMS
from typesafe_sdk import Choice, TypeSafeClient
JEV_MODEL = "jev-1.13.0"
JEV_INPUT_USD_PER_MILLION_TOKENS = 0.042
CONFIDENCE_THRESHOLD = 0.70
CLAUDE_CALL_BUDGET_USD = 0.25
What do these constants control?
- The TypeSafe imports provide the fixed-option questions plus the synchronous client.
- The JEV_MODEL constant pins requests to jev-1.13.0.
- The input price constant lets the router estimate Jev cost at $0.042 per million input tokens.
- The CONFIDENCE_THRESHOLD value records the project policy threshold of 0.70.
- Save router.py.
- Confirm that the first line reads from __future__ import annotations.
Imports showing problems?
Check that the SDK import uses typesafe_sdk with an underscore. Keep the capitalization of Choice plus TypeSafeClient unchanged.
The policy import resolves after you create policy.py in the next substep.
Help me check the imports and constants in my router.
The router needs a consistent ordering for routes. It also needs a stable mapping from each route to the reviewer profile selected later.
- Add the route mappings directly below CLAUDE_CALL_BUDGET_USD = 0.25 by pasting this code:
ROUTE_RANK = {"routine": 0, "complex": 1, "high_risk": 2}
LEVEL_TO_ROUTE = {"low": "routine", "medium": "complex", "high": "high_risk"}
ROUTE_TO_PROFILE = {
"routine": "routine-reviewer",
"complex": "complex-reviewer",
"high_risk": "high-risk-reviewer",
}
PROFILE_TO_MODEL = {
"routine-reviewer": "haiku",
"complex-reviewer": "sonnet",
"high-risk-reviewer": "opus",
}
How do the mappings fit together?
Route ranks give Python an ordered scale. The router can select whichever signal implies the deeper review.
Profile mappings translate the chosen route into a reviewer name. Model mappings preserve the requested Claude model alias for later reporting.
- Save router.py.
- Confirm that high_risk has route rank 2.
Route mapping looks incomplete?
Check that every route appears in ROUTE_RANK plus ROUTE_TO_PROFILE.
Confirm that each opening brace has a matching closing brace.
Help me compare my route mappings with the expected router structure.
Typed records keep the model evidence together. The first record stores both answers plus the metadata needed for evaluation.
- Add the frozen JevSignals data class below the profile mappings by pasting this code:
@dataclass(frozen=True)
class JevSignals:
complexity: str
consequence: str
confidence: float
complexity_probabilities: dict[str, float]
consequence_probabilities: dict[str, float]
model: str
input_tokens: int
jev_cost_usd: float
What does JevSignals preserve?
The record stores the selected complexity plus consequence levels. It also keeps both probability maps.
Token usage supports cost calculation. The frozen data class prevents accidental changes after the evidence has been collected.
- Save router.py.
- Confirm that JevSignals contains nine fields.
Data class underlined?
Check that @dataclass(frozen=True) sits immediately above the class declaration.
Keep every field indented with four spaces.
Help me fix the JevSignals data class in router.py.
A route decision needs the selected tier plus the evidence that produced it. Its JSON conversion makes the complete decision inspectable from the terminal.
- Add the frozen RouteDecision data class below JevSignals by pasting this code:
@dataclass(frozen=True)
class RouteDecision:
route: str
profile: str
hard_gates: tuple[str, ...]
reasons: tuple[str, ...]
signals: JevSignals
def to_dict(self) -> dict[str, object]:
return {
"route": self.route,
"profile": self.profile,
"hard_gates": list(self.hard_gates),
"reasons": list(self.reasons),
"signals": asdict(self.signals),
}
Why convert the decision to a dictionary?
JSON can serialize dictionaries plus lists directly. The method converts tuple fields into lists while asdict() expands the nested signals record.
The terminal output can therefore show the route plus every reason behind it.
- Save router.py.
- Confirm that to_dict() returns five top-level keys.
Dictionary structure looks uneven?
Check that to_dict() remains indented inside RouteDecision.
Confirm that the closing brace aligns with the return statement.
Help me fix the RouteDecision data class and its to_dict method.
The model call evaluates both questions against the same unified diff. Each question stays focused on one dimension of review depth.
- Add the first half of get_jev_signals() below RouteDecision by pasting this code:
def get_jev_signals(diff: str) -> JevSignals:
questions = {
"complexity": Choice(
instructions="How difficult is this code change to reason about and verify?",
criteria={
"low": "Localized, mechanical, or documentation-only change with simple behavior.",
"medium": "Multiple branches, state transitions, integrations, or nontrivial tests.",
"high": "Cross-cutting architecture, concurrency, migrations, or difficult failure modes.",
},
),
"consequence": Choice(
instructions="What is the consequence if this change is wrong in production?",
criteria={
"low": "Minor developer experience or presentation impact with easy recovery.",
"medium": "User-visible incorrect behavior, degraded reliability, or operational toil.",
"high": "Security exposure, privilege failure, data loss, financial impact, or broad outage.",
},
),
}
with TypeSafeClient() as client:
response = client.system_one(
state={
"task": "Route this unified code diff to an appropriate review depth.",
"diff": diff,
},
questions=questions,
model=JEV_MODEL,
)
What happens in the model request?
- The complexity question measures how difficult the implementation is to reason about.
- The consequence question measures the impact of a production mistake.
- Both Choice objects use the same low, medium, plus high scale.
- The client sends one shared state object to the pinned Jev model.
- Save router.py.
- Confirm that the questions dictionary contains complexity plus consequence.
Choice criteria not lining up?
Check that each Choice has its own criteria dictionary. Each dictionary needs the same three level names.
Confirm that model=JEV_MODEL sits inside the system_one() call.
Help me check the two Choice questions and the TypeSafe request.
The response contains one result per question. The router uses the lower confidence because the least certain signal is the safest basis for escalation.
- Complete get_jev_signals() by adding this code directly below the system_one() call:
complexity = response.choices["complexity"]
consequence = response.choices["consequence"]
confidence = min(complexity.confidence, consequence.confidence)
input_tokens = response.usage.input_tokens or 0
jev_cost_usd = input_tokens * JEV_INPUT_USD_PER_MILLION_TOKENS / 1_000_000
return JevSignals(
complexity=complexity.choice,
consequence=consequence.choice,
confidence=confidence,
complexity_probabilities=dict(complexity.probabilities),
consequence_probabilities=dict(consequence.probabilities),
model=response.model,
input_tokens=input_tokens,
jev_cost_usd=jev_cost_usd,
)
How is the response normalized?
The minimum confidence makes uncertainty conservative. Either uncertain answer can trigger deeper review.
The router retains both complete probability maps. It also converts input tokens into an estimated Jev cost.
- Save router.py.
- Confirm that the function returns one JevSignals object.
Response fields showing syntax problems?
Keep this entire chunk indented inside get_jev_signals().
Check that both probability fields wrap their source values with dict().
Help me check the response handling in get_jev_signals.
Convert evidence into deterministic policy
Model signals describe a change. A safety gate encodes categories that always require the deepest route.
The policy first chooses the highest tier implied by either model answer. Low confidence raises that result by one tier.
- Select the new-file control in the Explorer sidebar.
- Name the new file policy.py.
- Add the initial hard-gate terms by pasting this code:
HARD_GATE_TERMS = (
"api key",
"credential",
"password",
"secret",
"payment",
"billing",
"charge_card",
"drop table",
"delete from",
"truncate table",
)
What does this policy protect?
These terms cover credentials, payments, plus destructive database operations. A matching diff receives the high-risk route regardless of the model result.
The tuple keeps non-negotiable rules in ordinary source code. Reviewers can inspect every protected category.
- Save policy.py.
- Confirm that policy.py appears beside router.py in the Explorer sidebar.
Policy import still unresolved?
Check that policy.py sits in the same project folder as router.py.
Confirm that the exported name is exactly HARD_GATE_TERMS.
Help me resolve the policy import in router.py.
✔️ Awesome, I've got everything!
Your initial safety terms are now stored in policy.py.
ⓧ I'd like to double check the full code
HARD_GATE_TERMS = (
"api key",
"credential",
"password",
"secret",
"payment",
"billing",
"charge_card",
"drop table",
"delete from",
"truncate table",
)
How to compare policy.py
The reference tab shows the complete file at this point. Match every term plus the final closing parenthesis.
The route calculation is intentionally simple. It takes the deeper of the two model-derived routes before applying confidence escalation plus hard gates.
- Return to router.py in the editor.
- Find the line def get_jev_signals(diff: str) -> JevSignals:.
- Insert the first half of the policy logic immediately above that line by pasting this code:
def _higher_route(route: str) -> str:
return {"routine": "complex", "complex": "high_risk", "high_risk": "high_risk"}[route]
def apply_policy(diff: str, signals: JevSignals) -> RouteDecision:
complexity_route = LEVEL_TO_ROUTE[signals.complexity]
consequence_route = LEVEL_TO_ROUTE[signals.consequence]
route = max((complexity_route, consequence_route), key=ROUTE_RANK.get)
reasons = [
f"complexity={signals.complexity}",
f"consequence={signals.consequence}",
]
How is the starting route chosen?
Each level first becomes a named route. Python then uses route ranks to keep whichever route is deeper.
The reasons list preserves both source signals. Every decision can explain its starting point.
- Save router.py.
- Confirm that _higher_route() leaves high_risk at the same tier.
Route comparison underlined?
Check that ROUTE_RANK.get is passed as the key argument to max().
Keep the reasons list indented inside apply_policy().
Help me fix the initial route calculation in apply_policy.
Confidence escalation handles uncertain judgments. Hard gates handle known categories where a downgrade is unacceptable.
- Complete apply_policy() by adding this code directly below the reasons list:
if signals.confidence < CONFIDENCE_THRESHOLD:
route = _higher_route(route)
reasons.append(
f"confidence={signals.confidence:.2f} below {CONFIDENCE_THRESHOLD:.2f}; escalated one tier"
)
lowered = diff.lower()
hard_gates = tuple(term for term in HARD_GATE_TERMS if term in lowered)
if hard_gates:
route = "high_risk"
reasons.append("non-negotiable safety gate matched")
return RouteDecision(
route=route,
profile=ROUTE_TO_PROFILE[route],
hard_gates=hard_gates,
reasons=tuple(reasons),
signals=signals,
)
How does the policy fail safely?
Confidence below the threshold raises the selected route by one level. A high-risk route remains high-risk.
The diff is normalized to lowercase before term matching. Any hard-gate match forces the final route to high_risk.
- Save router.py.
- Confirm that apply_policy() returns a RouteDecision.
Policy function ending too early?
Keep the confidence check plus hard-gate check inside apply_policy().
Align the final return RouteDecision( with the two if statements.
Help me fix the confidence and hard-gate logic in apply_policy.
Add and test the dry-run path
A dry run stops after routing. This lets you inspect Jev evidence plus Python policy before any reviewer receives the diff.
This step exercises only the --dry-run branch. The review branch names dispatch_review() for the reviewer-dispatch step that follows.
- Add the route wrapper plus command-line entry point below get_jev_signals() by pasting this code:
def route_change(diff: str) -> RouteDecision:
return apply_policy(diff, get_jev_signals(diff))
def main() -> None:
parser = argparse.ArgumentParser(description="Route and review a unified code diff.")
parser.add_argument("diff_file", nargs="?", type=Path)
parser.add_argument("--dry-run", action="store_true")
args = parser.parse_args()
diff = args.diff_file.read_text() if args.diff_file else sys.stdin.read()
if not diff.strip():
raise SystemExit("Provide a diff file or pipe a diff on stdin.")
decision = route_change(diff)
output: dict[str, object] = {"decision": decision.to_dict()}
if not args.dry_run:
output["review"] = dispatch_review(diff, decision)
print(json.dumps(output, indent=2))
if __name__ == "__main__":
main()
What does the command-line path do?
The router reads a named diff file when one is provided. Otherwise it reads piped content from standard input.
The dry-run flag leaves the output focused on the decision. Indented JSON makes every signal plus policy reason readable.
- Save router.py.
- Confirm that the final line calls main().
Command-line block showing indentation problems?
Check that argument parsing plus input handling remain inside main().
Keep the final if __name__ == "__main__": block aligned at the far left.
Help me fix the route wrapper and command-line entry point.
✔️ Awesome, I've got everything!
Your router.py file now contains the typed Jev decision path plus deterministic policy.
ⓧ I'd like to double check the full code
from __future__ import annotations
import argparse
import json
import subprocess
import sys
from dataclasses import asdict, dataclass
from pathlib import Path
from policy import HARD_GATE_TERMS
from typesafe_sdk import Choice, TypeSafeClient
JEV_MODEL = "jev-1.13.0"
JEV_INPUT_USD_PER_MILLION_TOKENS = 0.042
CONFIDENCE_THRESHOLD = 0.70
CLAUDE_CALL_BUDGET_USD = 0.25
ROUTE_RANK = {"routine": 0, "complex": 1, "high_risk": 2}
LEVEL_TO_ROUTE = {"low": "routine", "medium": "complex", "high": "high_risk"}
ROUTE_TO_PROFILE = {
"routine": "routine-reviewer",
"complex": "complex-reviewer",
"high_risk": "high-risk-reviewer",
}
PROFILE_TO_MODEL = {
"routine-reviewer": "haiku",
"complex-reviewer": "sonnet",
"high-risk-reviewer": "opus",
}
@dataclass(frozen=True)
class JevSignals:
complexity: str
consequence: str
confidence: float
complexity_probabilities: dict[str, float]
consequence_probabilities: dict[str, float]
model: str
input_tokens: int
jev_cost_usd: float
@dataclass(frozen=True)
class RouteDecision:
route: str
profile: str
hard_gates: tuple[str, ...]
reasons: tuple[str, ...]
signals: JevSignals
def to_dict(self) -> dict[str, object]:
return {
"route": self.route,
"profile": self.profile,
"hard_gates": list(self.hard_gates),
"reasons": list(self.reasons),
"signals": asdict(self.signals),
}
def _higher_route(route: str) -> str:
return {"routine": "complex", "complex": "high_risk", "high_risk": "high_risk"}[route]
def apply_policy(diff: str, signals: JevSignals) -> RouteDecision:
complexity_route = LEVEL_TO_ROUTE[signals.complexity]
consequence_route = LEVEL_TO_ROUTE[signals.consequence]
route = max((complexity_route, consequence_route), key=ROUTE_RANK.get)
reasons = [
f"complexity={signals.complexity}",
f"consequence={signals.consequence}",
]
if signals.confidence < CONFIDENCE_THRESHOLD:
route = _higher_route(route)
reasons.append(
f"confidence={signals.confidence:.2f} below {CONFIDENCE_THRESHOLD:.2f}; escalated one tier"
)
lowered = diff.lower()
hard_gates = tuple(term for term in HARD_GATE_TERMS if term in lowered)
if hard_gates:
route = "high_risk"
reasons.append("non-negotiable safety gate matched")
return RouteDecision(
route=route,
profile=ROUTE_TO_PROFILE[route],
hard_gates=hard_gates,
reasons=tuple(reasons),
signals=signals,
)
def get_jev_signals(diff: str) -> JevSignals:
questions = {
"complexity": Choice(
instructions="How difficult is this code change to reason about and verify?",
criteria={
"low": "Localized, mechanical, or documentation-only change with simple behavior.",
"medium": "Multiple branches, state transitions, integrations, or nontrivial tests.",
"high": "Cross-cutting architecture, concurrency, migrations, or difficult failure modes.",
},
),
"consequence": Choice(
instructions="What is the consequence if this change is wrong in production?",
criteria={
"low": "Minor developer experience or presentation impact with easy recovery.",
"medium": "User-visible incorrect behavior, degraded reliability, or operational toil.",
"high": "Security exposure, privilege failure, data loss, financial impact, or broad outage.",
},
),
}
with TypeSafeClient() as client:
response = client.system_one(
state={
"task": "Route this unified code diff to an appropriate review depth.",
"diff": diff,
},
questions=questions,
model=JEV_MODEL,
)
complexity = response.choices["complexity"]
consequence = response.choices["consequence"]
confidence = min(complexity.confidence, consequence.confidence)
input_tokens = response.usage.input_tokens or 0
jev_cost_usd = input_tokens * JEV_INPUT_USD_PER_MILLION_TOKENS / 1_000_000
return JevSignals(
complexity=complexity.choice,
consequence=consequence.choice,
confidence=confidence,
complexity_probabilities=dict(complexity.probabilities),
consequence_probabilities=dict(consequence.probabilities),
model=response.model,
input_tokens=input_tokens,
jev_cost_usd=jev_cost_usd,
)
def route_change(diff: str) -> RouteDecision:
return apply_policy(diff, get_jev_signals(diff))
def main() -> None:
parser = argparse.ArgumentParser(description="Route and review a unified code diff.")
parser.add_argument("diff_file", nargs="?", type=Path)
parser.add_argument("--dry-run", action="store_true")
args = parser.parse_args()
diff = args.diff_file.read_text() if args.diff_file else sys.stdin.read()
if not diff.strip():
raise SystemExit("Provide a diff file or pipe a diff on stdin.")
decision = route_change(diff)
output: dict[str, object] = {"decision": decision.to_dict()}
if not args.dry_run:
output["review"] = dispatch_review(diff, decision)
print(json.dumps(output, indent=2))
if __name__ == "__main__":
main()
How to compare router.py
Compare the reference from top to bottom. Pay close attention to indentation around both data classes plus the three functions.
The dry-run path is complete at this point. The later reviewer step supplies the dispatcher used by the non-dry branch.
Before you test the router, predict which evidence fields will appear for a harmless documentation change.
- Run the harmless documentation diff through the dry-run path with this command:
printf '%s\n' 'diff --git a/README.md b/README.md' '-Run the server.' '+Run the local server.' | python router.py --dry-run
What did the dry run prove?
The JSON contains complexity, consequence, plus confidence inside the signals object. Separate probability maps preserve the full Choice results.
The decision also contains route, profile, plus policy reasons. No Claude Code review starts during this dry run.
That closes the routing loop. Your harmless diff now produces typed evidence plus an inspectable Python-owned decision.
Dry run not producing decision JSON?
Confirm that the virtual environment from the setup step remains active. Check that TYPESAFE_API_KEY is still available in the current terminal.
Compare the indentation in get_jev_signals() plus apply_policy() with the full-code tab.
Help me diagnose why the router dry run is not returning decision JSON.
Your router can now turn model evidence into a deterministic review tier. Next up, you will connect those tiers to progressive read-only Claude Code reviewers.
Dispatch Progressive Reviewers
Your Jev router now turns typed evidence into an inspectable review route. Confidence escalation already gives uncertain changes more scrutiny.
A route becomes useful when it selects a matching Claude Code reviewer. Deterministic safety gates keep payment changes in the deepest review tier.
In this step, get ready to:
- Create three read-only reviewer profiles with progressively deeper review instructions.
- Connect live Claude Code dispatch to the selected route.
- Prove that a payment hard gate forces the high-risk reviewer.
Create progressive reviewer profiles
Project-scoped agents live in .claude/agents/. Each profile requests a model alias that matches its intended review depth.
Why use model aliases?
The aliases haiku, sonnet, and opus express progressively deeper review levels. Your organization can substitute an allowed model when an alias is unavailable.
The live response records the resolved model names. That keeps requested depth separate from actual model usage.
- In the Visual Studio Code Explorer sidebar, select the New Folder icon.
- Enter .claude as the folder name.
- Select the .claude folder.
- Select the New Folder icon.
- Enter agents as the nested folder name.
The .claude/agents/ folder gives Claude Code a project-local catalogue of reviewer profiles.
- Select the agents folder in the Explorer sidebar.
- Select the New File icon.
- Enter routine-reviewer.md as the file name.
- Add the routine reviewer profile by pasting this content:
---
name: routine-reviewer
description: Reviews low-complexity and low-consequence diffs with a concise correctness pass.
model: haiku
effort: low
---
You are a concise, read-only code reviewer for routine changes.
Check only for direct correctness errors, misleading wording, obvious regressions, and missing focused tests. Avoid speculative architecture advice.
Return a verdict, concrete findings ordered by severity, and at most two targeted checks.
How does the routine reviewer work?
- The frontmatter names the routine-reviewer profile.
- The haiku alias requests the lightest review model in this project.
- The instructions focus the response on direct defects plus targeted checks.
- Save .claude/agents/routine-reviewer.md.
- Confirm the file appears beneath .claude/agents/ in the Explorer sidebar.
The routine profile is ready. Low-complexity changes now have a concise review path.
Routine profile missing?
- Confirm that routine-reviewer.md sits inside .claude/agents/.
- Check that the file name ends with .md.
Ask for help with checking the routine reviewer file location and frontmatter.
- Select the agents folder in the Explorer sidebar.
- Select the New File icon.
- Enter complex-reviewer.md as the file name.
- Add the complex reviewer profile by pasting this content:
---
name: complex-reviewer
description: Reviews changes with branching, state, integrations, or cross-file behavior.
model: sonnet
effort: high
---
You are a senior, read-only code reviewer for complex changes.
Trace control flow, state transitions, error handling, compatibility, failure recovery, and test coverage. Distinguish demonstrated defects from questions.
Return a verdict, critical and major findings with evidence, edge cases, targeted tests, and residual uncertainty.
How does the complex reviewer work?
- The complex-reviewer profile requests the sonnet alias.
- Its instructions trace control flow plus failure recovery.
- The response separates demonstrated defects from unresolved questions.
- Save .claude/agents/complex-reviewer.md.
- Confirm the file appears beside routine-reviewer.md in the Explorer sidebar.
The complex profile now covers branching plus stateful behavior. Its review instructions demand deeper evidence.
Complex profile missing?
- Confirm that complex-reviewer.md sits beside routine-reviewer.md.
- Check that model: sonnet remains inside the frontmatter.
Ask for help with validating the complex reviewer profile.
- Select the agents folder in the Explorer sidebar.
- Select the New File icon.
- Enter high-risk-reviewer.md as the file name.
- Add the high-risk reviewer profile by pasting this content:
---
name: high-risk-reviewer
description: Performs deep review for security, authorization, credentials, payments, data loss, and destructive operations.
model: opus
effort: high
---
You are a read-only security and consequence reviewer for high-risk changes.
Assume the diff may cross a trust boundary. Analyze authorization, authentication, secret handling, payment correctness, destructive data operations, bypass paths, rollback, auditability, and fail-open behavior. Treat small diffs as potentially high consequence.
Return a verdict, exploit or failure scenarios, findings ordered by severity with exact diff evidence, required controls before merge, and residual risk.
How does the high-risk reviewer work?
- The high-risk-reviewer profile requests the opus alias.
- Its review scope includes trust boundaries plus destructive operations.
- Its output requires concrete failure scenarios plus controls before merge.
- Save .claude/agents/high-risk-reviewer.md.
- Confirm all three reviewer files appear beneath .claude/agents/.
You now have three project-scoped reviewers. Each route has a read-only profile with a matching depth.
High-risk profile missing?
- Confirm that the file name uses high-risk-reviewer.md exactly.
- Check that model: opus remains inside the frontmatter.
Ask for help with validating the high-risk reviewer profile.
Connect live dispatch to the router
The selected profile now needs a safe execution path. dispatch_review() builds a constrained prompt before invoking Claude Code in non-interactive mode.
- Return to router.py in Visual Studio Code.
- Place your cursor after route_change().
- Add the first part of dispatch_review() by pasting this code:
def dispatch_review(diff: str, decision: RouteDecision) -> dict[str, object]:
prompt = f"""Review only the supplied unified diff.
Routing evidence:
- route: {decision.route}
- complexity: {decision.signals.complexity}
- consequence: {decision.signals.consequence}
- confidence: {decision.signals.confidence:.2f}
- hard gates: {', '.join(decision.hard_gates) or 'none'}
Report concrete findings with severity, evidence from the diff, and a suggested check.
Do not edit files or run tools.
Unified diff:
{diff}
"""
What does the review prompt do?
- The prompt includes the route plus each Jev signal.
- The hard-gate list shows which deterministic rule influenced the route.
- The final instruction keeps the reviewer focused on evidence from the supplied diff.
- Save router.py.
- Open the Visual Studio Code Outline view.
- Confirm that dispatch_review() appears beneath route_change().
The new function now has a visible prompt-building entry point in your router.
Dispatch function missing from Outline?
- Check that def dispatch_review begins at the left edge of router.py.
- Confirm that the closing triple quotes align with the prompt assignment.
Ask for help with finding the indentation problem in dispatch_review().
- Inside dispatch_review(), place your cursor after the closing triple quotes.
- Add the constrained Claude Code command plus subprocess call by pasting this code:
command = [
"claude",
"-p",
"--agent",
decision.profile,
"--output-format",
"json",
"--tools",
"",
"--max-turns",
"1",
"--max-budget-usd",
str(CLAUDE_CALL_BUDGET_USD),
"--no-session-persistence",
prompt,
]
completed = subprocess.run(
command,
capture_output=True,
text=True,
timeout=180,
check=False,
)
How is the review constrained?
- The selected profile is passed through --agent.
- The empty --tools value disables built-in tools.
- The command allows one turn with no session persistence.
- The client-side estimate is capped by CLAUDE_CALL_BUDGET_USD.
- Save router.py.
- Open the Visual Studio Code Problems panel.
- Confirm that the new command section has no Python syntax errors.
The router can now launch one constrained reviewer process with the chosen profile.
Seeing a Python syntax problem?
- Confirm that every command value remains inside the command list.
- Check that subprocess.run() aligns with the command assignment.
Ask for help with checking the command list and subprocess indentation.
- Inside dispatch_review(), place your cursor after the subprocess.run() call.
- Add failure handling plus structured review metadata by pasting this code:
if completed.returncode != 0:
detail = completed.stderr.strip() or completed.stdout.strip()
raise RuntimeError(f"Claude Code failed with exit {completed.returncode}: {detail}")
payload = json.loads(completed.stdout)
return {
"profile": decision.profile,
"requested_model": PROFILE_TO_MODEL[decision.profile],
"result": payload.get("result", ""),
"claude_cost_usd": float(payload.get("total_cost_usd") or 0.0),
"model_usage": payload.get("modelUsage", {}),
}
What comes back from dispatch?
- A nonzero exit becomes a Python error with captured command details.
- The returned dictionary records the selected profile plus requested alias.
- The live payload contributes review text plus resolved model usage.
- The claude_cost_usd field stores the client-side cost estimate.
- Save router.py.
- Check the Visual Studio Code Problems panel for Python syntax errors.
The dispatch function now converts Claude Code JSON into the telemetry your evaluator can aggregate.
Return dictionary looks incomplete?
- Confirm that the dictionary contains profile, requested_model, result, claude_cost_usd, and model_usage.
- Check that the return dictionary remains indented inside dispatch_review().
Ask for help with checking the dispatch response fields.
The command path still needs to call the dispatcher. Dry runs keep returning the decision alone.
- Select the entire existing main() function in router.py.
- Replace that function with this version:
def main() -> None:
parser = argparse.ArgumentParser(description="Route and review a unified code diff.")
parser.add_argument("diff_file", nargs="?", type=Path)
parser.add_argument("--dry-run", action="store_true")
args = parser.parse_args()
diff = args.diff_file.read_text() if args.diff_file else sys.stdin.read()
if not diff.strip():
raise SystemExit("Provide a diff file or pipe a diff on stdin.")
decision = route_change(diff)
output: dict[str, object] = {"decision": decision.to_dict()}
if not args.dry_run:
output["review"] = dispatch_review(diff, decision)
print(json.dumps(output, indent=2))
How does the main path change?
- Every input still passes through route_change() first.
- The --dry-run path skips the live review call.
- The live path adds a review object beside the decision.
- Save router.py.
- Confirm that the Visual Studio Code Problems panel remains clear of Python syntax errors.
The dispatch path is connected. A live run now produces the route plus the selected reviewer's response metadata.
Live path still missing?
- Confirm that output["review"] sits inside the if not args.dry_run: block.
- Check that the final print() remains outside that conditional block.
Ask for help with checking the main path indentation.
Double-check the complete project code
The full reference below shows the cumulative project state for this step. Use it to compare file names plus exact contents.
✔️ Awesome, I've got everything!
Great. Save every file before running the live high-risk check.
ⓧ I'd like to double check the full code
Compare each project file with the matching reference below.
typesafe-sdk==0.7.2
What belongs in requirements.txt?
This file keeps the TypeSafe SDK dependency pinned to the project version.
HARD_GATE_TERMS = (
"api key",
"credential",
"password",
"secret",
"payment",
"billing",
"charge_card",
"drop table",
"delete from",
"truncate table",
)
What belongs in policy.py?
These initial hard-gate terms force credentials plus payment or destructive database changes into the high-risk route.
from __future__ import annotations
import json
from pathlib import Path
ROUTE_RANK = {"routine": 0, "complex": 1, "high_risk": 2}
def main() -> None:
cases = json.loads(Path("eval_cases.json").read_text())
predicted = "complex"
correct = 0
over = 0
under = 0
print("case\texpected\tpredicted")
for case in cases:
expected = case["expected_route"]
correct += predicted == expected
over += ROUTE_RANK[predicted] > ROUTE_RANK[expected]
under += ROUTE_RANK[predicted] < ROUTE_RANK[expected]
print(f"{case['id']}\t{expected}\t{predicted}")
print(f"\naccuracy={correct / len(cases):.1%}")
print(f"over_escalations={over}")
print(f"under_escalations={under}")
if __name__ == "__main__":
main()
What belongs in baseline.py?
The baseline keeps the fixed complex prediction used to expose over-escalation plus under-escalation.
[
{"id":"docs-wording","expected_route":"routine","diff":"diff --git a/README.md b/README.md\n-Run the server.\n+Run the local development server.\n"},
{"id":"local-rename","expected_route":"routine","diff":"diff --git a/items.py b/items.py\n-result = [item for item in items if item.active]\n-return result\n+active_items = [item for item in items if item.active]\n+return active_items\n"},
{"id":"retry-loop","expected_route":"complex","diff":"diff --git a/client.py b/client.py\n+for attempt in range(3):\n+ try:\n+ return send_request()\n+ except TimeoutError:\n+ if attempt == 2:\n+ raise\n"},
{"id":"schema-extension","expected_route":"complex","diff":"diff --git a/migrations/014_orders.sql b/migrations/014_orders.sql\n+ALTER TABLE orders ADD COLUMN external_id TEXT;\n+CREATE INDEX orders_external_id_idx ON orders(external_id);\n"},
{"id":"payment-boundary","expected_route":"high_risk","diff":"diff --git a/payments.py b/payments.py\n-if amount > 0:\n+if amount >= 0:\n charge_card(amount)\n"},
{"id":"destructive-migration","expected_route":"high_risk","diff":"diff --git a/migrations/015_cleanup.sql b/migrations/015_cleanup.sql\n+DROP TABLE archived_orders;\n"},
{"id":"three-line-authorization","expected_route":"high_risk","diff":"diff --git a/authorization.py b/authorization.py\n-if user.is_admin:\n+if user:\n return view_audit_log()\n"}
]
What belongs in eval_cases.json?
The seven labeled diffs cover routine changes plus complex behavior plus high-consequence cases.
---
name: routine-reviewer
description: Reviews low-complexity and low-consequence diffs with a concise correctness pass.
model: haiku
effort: low
---
You are a concise, read-only code reviewer for routine changes.
Check only for direct correctness errors, misleading wording, obvious regressions, and missing focused tests. Avoid speculative architecture advice.
Return a verdict, concrete findings ordered by severity, and at most two targeted checks.
What belongs in the routine profile?
This profile requests Haiku for concise review of routine changes.
---
name: complex-reviewer
description: Reviews changes with branching, state, integrations, or cross-file behavior.
model: sonnet
effort: high
---
You are a senior, read-only code reviewer for complex changes.
Trace control flow, state transitions, error handling, compatibility, failure recovery, and test coverage. Distinguish demonstrated defects from questions.
Return a verdict, critical and major findings with evidence, edge cases, targeted tests, and residual uncertainty.
What belongs in the complex profile?
This profile requests Sonnet for stateful or cross-file review.
---
name: high-risk-reviewer
description: Performs deep review for security, authorization, credentials, payments, data loss, and destructive operations.
model: opus
effort: high
---
You are a read-only security and consequence reviewer for high-risk changes.
Assume the diff may cross a trust boundary. Analyze authorization, authentication, secret handling, payment correctness, destructive data operations, bypass paths, rollback, auditability, and fail-open behavior. Treat small diffs as potentially high consequence.
Return a verdict, exploit or failure scenarios, findings ordered by severity with exact diff evidence, required controls before merge, and residual risk.
What belongs in the high-risk profile?
This profile requests Opus for deep consequence plus security review.
from __future__ import annotations
import argparse
import json
import subprocess
import sys
from dataclasses import asdict, dataclass
from pathlib import Path
from policy import HARD_GATE_TERMS
from typesafe_sdk import Choice, TypeSafeClient
JEV_MODEL = "jev-1.13.0"
JEV_INPUT_USD_PER_MILLION_TOKENS = 0.042
CONFIDENCE_THRESHOLD = 0.70
CLAUDE_CALL_BUDGET_USD = 0.25
ROUTE_RANK = {"routine": 0, "complex": 1, "high_risk": 2}
LEVEL_TO_ROUTE = {"low": "routine", "medium": "complex", "high": "high_risk"}
ROUTE_TO_PROFILE = {
"routine": "routine-reviewer",
"complex": "complex-reviewer",
"high_risk": "high-risk-reviewer",
}
PROFILE_TO_MODEL = {
"routine-reviewer": "haiku",
"complex-reviewer": "sonnet",
"high-risk-reviewer": "opus",
}
@dataclass(frozen=True)
class JevSignals:
complexity: str
consequence: str
confidence: float
complexity_probabilities: dict[str, float]
consequence_probabilities: dict[str, float]
model: str
input_tokens: int
jev_cost_usd: float
@dataclass(frozen=True)
class RouteDecision:
route: str
profile: str
hard_gates: tuple[str, ...]
reasons: tuple[str, ...]
signals: JevSignals
def to_dict(self) -> dict[str, object]:
return {
"route": self.route,
"profile": self.profile,
"hard_gates": list(self.hard_gates),
"reasons": list(self.reasons),
"signals": asdict(self.signals),
}
def _higher_route(route: str) -> str:
return {"routine": "complex", "complex": "high_risk", "high_risk": "high_risk"}[route]
def apply_policy(diff: str, signals: JevSignals) -> RouteDecision:
complexity_route = LEVEL_TO_ROUTE[signals.complexity]
consequence_route = LEVEL_TO_ROUTE[signals.consequence]
route = max((complexity_route, consequence_route), key=ROUTE_RANK.get)
reasons = [
f"complexity={signals.complexity}",
f"consequence={signals.consequence}",
]
if signals.confidence < CONFIDENCE_THRESHOLD:
route = _higher_route(route)
reasons.append(
f"confidence={signals.confidence:.2f} below {CONFIDENCE_THRESHOLD:.2f}; escalated one tier"
)
lowered = diff.lower()
hard_gates = tuple(term for term in HARD_GATE_TERMS if term in lowered)
if hard_gates:
route = "high_risk"
reasons.append("non-negotiable safety gate matched")
return RouteDecision(
route=route,
profile=ROUTE_TO_PROFILE[route],
hard_gates=hard_gates,
reasons=tuple(reasons),
signals=signals,
)
def get_jev_signals(diff: str) -> JevSignals:
questions = {
"complexity": Choice(
instructions="How difficult is this code change to reason about and verify?",
criteria={
"low": "Localized, mechanical, or documentation-only change with simple behavior.",
"medium": "Multiple branches, state transitions, integrations, or nontrivial tests.",
"high": "Cross-cutting architecture, concurrency, migrations, or difficult failure modes.",
},
),
"consequence": Choice(
instructions="What is the consequence if this change is wrong in production?",
criteria={
"low": "Minor developer experience or presentation impact with easy recovery.",
"medium": "User-visible incorrect behavior, degraded reliability, or operational toil.",
"high": "Security exposure, privilege failure, data loss, financial impact, or broad outage.",
},
),
}
with TypeSafeClient() as client:
response = client.system_one(
state={
"task": "Route this unified code diff to an appropriate review depth.",
"diff": diff,
},
questions=questions,
model=JEV_MODEL,
)
complexity = response.choices["complexity"]
consequence = response.choices["consequence"]
confidence = min(complexity.confidence, consequence.confidence)
input_tokens = response.usage.input_tokens or 0
jev_cost_usd = input_tokens * JEV_INPUT_USD_PER_MILLION_TOKENS / 1_000_000
return JevSignals(
complexity=complexity.choice,
consequence=consequence.choice,
confidence=confidence,
complexity_probabilities=dict(complexity.probabilities),
consequence_probabilities=dict(consequence.probabilities),
model=response.model,
input_tokens=input_tokens,
jev_cost_usd=jev_cost_usd,
)
def route_change(diff: str) -> RouteDecision:
return apply_policy(diff, get_jev_signals(diff))
def dispatch_review(diff: str, decision: RouteDecision) -> dict[str, object]:
prompt = f"""Review only the supplied unified diff.
Routing evidence:
- route: {decision.route}
- complexity: {decision.signals.complexity}
- consequence: {decision.signals.consequence}
- confidence: {decision.signals.confidence:.2f}
- hard gates: {', '.join(decision.hard_gates) or 'none'}
Report concrete findings with severity, evidence from the diff, and a suggested check.
Do not edit files or run tools.
Unified diff:
{diff}
"""
command = [
"claude",
"-p",
"--agent",
decision.profile,
"--output-format",
"json",
"--tools",
"",
"--max-turns",
"1",
"--max-budget-usd",
str(CLAUDE_CALL_BUDGET_USD),
"--no-session-persistence",
prompt,
]
completed = subprocess.run(
command,
capture_output=True,
text=True,
timeout=180,
check=False,
)
if completed.returncode != 0:
detail = completed.stderr.strip() or completed.stdout.strip()
raise RuntimeError(f"Claude Code failed with exit {completed.returncode}: {detail}")
payload = json.loads(completed.stdout)
return {
"profile": decision.profile,
"requested_model": PROFILE_TO_MODEL[decision.profile],
"result": payload.get("result", ""),
"claude_cost_usd": float(payload.get("total_cost_usd") or 0.0),
"model_usage": payload.get("modelUsage", {}),
}
def main() -> None:
parser = argparse.ArgumentParser(description="Route and review a unified code diff.")
parser.add_argument("diff_file", nargs="?", type=Path)
parser.add_argument("--dry-run", action="store_true")
args = parser.parse_args()
diff = args.diff_file.read_text() if args.diff_file else sys.stdin.read()
if not diff.strip():
raise SystemExit("Provide a diff file or pipe a diff on stdin.")
decision = route_change(diff)
output: dict[str, object] = {"decision": decision.to_dict()}
if not args.dry_run:
output["review"] = dispatch_review(diff, decision)
print(json.dumps(output, indent=2))
if __name__ == "__main__":
main()
What belongs in router.py?
This cumulative file keeps Jev evidence plus deterministic routing plus constrained reviewer dispatch in one terminal pipeline.
Test the guaranteed high-risk path
The current hard-gate list already includes payment terms plus charge_card. A matching term forces high_risk after the Jev decision.
Live review cost
This check makes one metered Jev request plus one live Claude Code review. The Claude Code client-side estimate for this call is capped at $0.25.
The estimate can differ from the final bill. One payment diff is enough to prove the dispatch path.
Before you run this, which profile do you expect the payment hard gate to select?
- Pipe the payment diff into the live router by running this command:
printf '%s\n' 'diff --git a/payments.py b/payments.py' '-if amount > 0:' '+if amount >= 0:' ' charge_card(amount)' | python router.py
What does this test prove?
- The diff contains payment-related terms covered by HARD_GATE_TERMS.
- The router still collects typed Jev signals before applying its hard gate.
- The selected reviewer runs with no built-in tools plus one allowed turn.
You should see high_risk as the route. The profile should be high-risk-reviewer.
The review object should contain requested_model, result, claude_cost_usd, and model_usage. The keys inside model_usage show the resolved model names reported by Claude Code.
That is the live routing loop working. Deterministic payment policy now selects the deepest read-only reviewer.
Live review did not complete?
- Confirm that your virtual environment remains active.
- Confirm that Claude Code remains authenticated.
- Check that high-risk-reviewer.md sits inside .claude/agents/.
Ask for help with diagnosing the live dispatch failure.
Your router now reserves progressively deeper reviewers for progressively riskier changes. Next, you will evaluate routing accuracy plus escalation errors plus model cost across all seven labeled cases.
Measure Routing Quality and Cost
Your router now selects a matching Claude Code reviewer for each diff. The hard gates also protect payment changes plus destructive database operations.
An adaptive workflow becomes trustworthy when labeled examples expose routing mistakes. This step measures those mistakes plus the model budget spent by each review path.
In this step, get ready to:
- Evaluate all seven labeled diffs through the adaptive router.
- Measure exact matches plus escalation mistakes.
- Write routing telemetry to reports/evaluation.json.
Build the labeled evaluator
A labeled evaluation compares each predicted route with the expected route stored in eval_cases.json. Route ranks reveal whether a mismatch used too much review depth or too little.
- Return to the Explorer sidebar in Visual Studio Code.
- Create evaluate.py inside the jev-review-router folder.
- Add the evaluator setup by pasting this code into evaluate.py:
from __future__ import annotations
import argparse
import json
from collections import Counter
from pathlib import Path
from router import PROFILE_TO_MODEL, ROUTE_RANK, dispatch_review, route_change
def main() -> None:
parser = argparse.ArgumentParser(description="Evaluate adaptive code review routing.")
parser.add_argument("--reviews", action="store_true", help="Run the selected Claude Code reviewer.")
args = parser.parse_args()
cases = json.loads(Path("eval_cases.json").read_text())
rows: list[dict[str, object]] = []
print("case\texpected\tpredicted\tprofile\tconfidence")
What Does This Code Set Up?
- The argument parser makes live reviews optional through --reviews.
- The cases variable loads the seven labeled diffs from eval_cases.json.
- The rows list keeps each decision available for the final report.
- The printed header gives every case a consistent set of comparison columns.
- Save evaluate.py.
- Confirm that evaluate.py appears beside router.py in the Explorer sidebar.
Cannot Find the Evaluator File?
Check that evaluate.py sits directly inside jev-review-router. A file created inside .claude cannot find the project files through the expected relative paths.
Help me place the evaluator correctly.
The first runnable slice should process every case before the summary logic is added. That gives you an immediate view of the routes produced by Jev plus deterministic policy.
- Place the case evaluation loop directly below the printed header by adding this code:
for case in cases:
decision = route_change(case["diff"])
row: dict[str, object] = {
"id": case["id"],
"expected_route": case["expected_route"],
"predicted_route": decision.route,
"profile": decision.profile,
"decision": decision.to_dict(),
}
if args.reviews:
row["review"] = dispatch_review(case["diff"], decision)
rows.append(row)
print(
f"{case['id']}\t{case['expected_route']}\t{decision.route}\t"
f"{decision.profile}\t{decision.signals.confidence:.2f}"
)
if __name__ == "__main__":
main()
How Does the Evaluation Loop Work?
- The loop sends each stored diff through route_change().
- Each row preserves the expected route plus the complete decision.
- The optional branch calls dispatch_review() only when live reviews are enabled.
- The final print displays the predicted profile plus the minimum Choice confidence for each case.
- Save evaluate.py.
Routing-Only Runs Still Use Jev
The routing-only check makes one metered Jev request for each of the seven cases. Jev 1.13 input costs $0.042 per million tokens.
This path skips every Claude Code reviewer call. It is the lower-cost way to inspect routing behavior.
Before you run this first pass, predict whether the seven labeled changes will all receive the same reviewer profile.
- Process the seven labeled cases by running:
python evaluate.py
What Should You See?
You should see seven rows beneath the case table header. Each row shows its expected route plus the predicted route plus the selected reviewer profile plus confidence.
The mix of profiles proves that the adaptive router can allocate different review depths across the same evaluation set.
Missing Cases or Routes?
Confirm that the active terminal remains inside jev-review-router. Check that eval_cases.json still contains all seven labeled cases.
Help me debug the incomplete evaluation.
Add quality and cost telemetry
A route mismatch has direction. An over-escalation spends more review depth than the label expects. An under-escalation gives a consequential change less scrutiny than expected.
The report also separates requested model aliases from resolved model names. This distinction preserves what the router requested while recording any model substitution applied by the learner's organization.
- In evaluate.py, locate the entry-point guard that begins with if __name__ == "__main__":.
- Insert the quality plus usage aggregation directly above that guard by adding:
correct = sum(row["expected_route"] == row["predicted_route"] for row in rows)
over = sum(
ROUTE_RANK[str(row["predicted_route"])] > ROUTE_RANK[str(row["expected_route"])]
for row in rows
)
under = sum(
ROUTE_RANK[str(row["predicted_route"])] < ROUTE_RANK[str(row["expected_route"])]
for row in rows
)
reviewer_usage = Counter(str(row["profile"]) for row in rows)
requested_model_usage = Counter(PROFILE_TO_MODEL[profile] for profile in reviewer_usage.elements())
resolved_model_usage: Counter[str] = Counter()
claude_cost = 0.0
for row in rows:
review = row.get("review")
if not isinstance(review, dict):
continue
claude_cost += float(review["claude_cost_usd"])
model_usage = review.get("model_usage", {})
if isinstance(model_usage, dict):
resolved_model_usage.update(model_usage.keys())
jev_input_tokens = sum(
int(row["decision"]["signals"]["input_tokens"]) # type: ignore[index]
for row in rows
)
What Do These Counters Measure?
- The correct count records exact route matches.
- The over plus under counts compare route ranks to identify the direction of each mismatch.
- The reviewer plus requested-model counters show the depth that the policy selected.
- The resolved-model counter reads the actual model names returned in modelUsage.
- The token sum gathers the metered Jev input from every decision.
- Save evaluate.py.
- Confirm that the new aggregation block remains indented inside main().
The final block turns the accumulated counters into a machine-readable summary. It also persists every case decision so you can inspect the evidence behind the totals.
- Continue directly below the jev_input_tokens calculation by adding:
jev_cost = sum(
float(row["decision"]["signals"]["jev_cost_usd"]) # type: ignore[index]
for row in rows
)
summary = {
"cases": len(rows),
"accuracy": correct / len(rows),
"over_escalations": over,
"under_escalations": under,
"reviewer_usage": dict(reviewer_usage),
"requested_model_usage": dict(requested_model_usage),
"resolved_model_usage": dict(resolved_model_usage),
"jev_input_tokens": jev_input_tokens,
"jev_cost_usd": jev_cost,
"claude_code_estimated_cost_usd": claude_cost,
"live_reviews_enabled": args.reviews,
}
report = {"summary": summary, "cases": rows}
reports = Path("reports")
reports.mkdir(exist_ok=True)
report_path = reports / "evaluation.json"
report_path.write_text(json.dumps(report, indent=2))
print("\nsummary")
print(json.dumps(summary, indent=2))
print(f"report={report_path}")
How Is the Report Produced?
- The jev_cost sum combines the input cost already calculated for every decision.
- The summary dictionary collects routing quality plus reviewer usage plus model usage plus cost telemetry.
- The report object keeps that summary beside every per-case decision.
- The file-writing block creates reports when needed. It writes formatted JSON to reports/evaluation.json.
✔️ Awesome, I've got everything!
Great. Save evaluate.py before running the complete routing-only evaluation.
ⓧ I'd like to double check the full code
from __future__ import annotations
import argparse
import json
from collections import Counter
from pathlib import Path
from router import PROFILE_TO_MODEL, ROUTE_RANK, dispatch_review, route_change
def main() -> None:
parser = argparse.ArgumentParser(description="Evaluate adaptive code review routing.")
parser.add_argument("--reviews", action="store_true", help="Run the selected Claude Code reviewer.")
args = parser.parse_args()
cases = json.loads(Path("eval_cases.json").read_text())
rows: list[dict[str, object]] = []
print("case\texpected\tpredicted\tprofile\tconfidence")
for case in cases:
decision = route_change(case["diff"])
row: dict[str, object] = {
"id": case["id"],
"expected_route": case["expected_route"],
"predicted_route": decision.route,
"profile": decision.profile,
"decision": decision.to_dict(),
}
if args.reviews:
row["review"] = dispatch_review(case["diff"], decision)
rows.append(row)
print(
f"{case['id']}\t{case['expected_route']}\t{decision.route}\t"
f"{decision.profile}\t{decision.signals.confidence:.2f}"
)
correct = sum(row["expected_route"] == row["predicted_route"] for row in rows)
over = sum(
ROUTE_RANK[str(row["predicted_route"])] > ROUTE_RANK[str(row["expected_route"])]
for row in rows
)
under = sum(
ROUTE_RANK[str(row["predicted_route"])] < ROUTE_RANK[str(row["expected_route"])]
for row in rows
)
reviewer_usage = Counter(str(row["profile"]) for row in rows)
requested_model_usage = Counter(PROFILE_TO_MODEL[profile] for profile in reviewer_usage.elements())
resolved_model_usage: Counter[str] = Counter()
claude_cost = 0.0
for row in rows:
review = row.get("review")
if not isinstance(review, dict):
continue
claude_cost += float(review["claude_cost_usd"])
model_usage = review.get("model_usage", {})
if isinstance(model_usage, dict):
resolved_model_usage.update(model_usage.keys())
jev_input_tokens = sum(
int(row["decision"]["signals"]["input_tokens"]) # type: ignore[index]
for row in rows
)
jev_cost = sum(
float(row["decision"]["signals"]["jev_cost_usd"]) # type: ignore[index]
for row in rows
)
summary = {
"cases": len(rows),
"accuracy": correct / len(rows),
"over_escalations": over,
"under_escalations": under,
"reviewer_usage": dict(reviewer_usage),
"requested_model_usage": dict(requested_model_usage),
"resolved_model_usage": dict(resolved_model_usage),
"jev_input_tokens": jev_input_tokens,
"jev_cost_usd": jev_cost,
"claude_code_estimated_cost_usd": claude_cost,
"live_reviews_enabled": args.reviews,
}
report = {"summary": summary, "cases": rows}
reports = Path("reports")
reports.mkdir(exist_ok=True)
report_path = reports / "evaluation.json"
report_path.write_text(json.dumps(report, indent=2))
print("\nsummary")
print(json.dumps(summary, indent=2))
print(f"report={report_path}")
if __name__ == "__main__":
main()
- Save evaluate.py.
Before you run the completed routing-only evaluation, predict which metric will reveal insufficient review depth.
- Generate the routing-only report by running:
python evaluate.py
What Should the Routing Report Show?
You should see the seven-case table followed by a formatted summary. The summary includes accuracy plus over-escalations plus under-escalations plus reviewer usage plus Jev cost.
You should also see report=reports/evaluation.json at the bottom. The generated report preserves every case decision while recording that live reviews were disabled.
- Confirm that reports/evaluation.json appears inside the new reports folder.
Report Missing or Summary Incomplete?
Check that the report-writing block remains inside main(). Confirm that it appears before the entry-point guard.
Help me debug the evaluation report.
Run the live evaluation
Routing-only evaluation measures policy behavior. Live evaluation also invokes the selected read-only reviewer once per case so the report can capture requested aliases plus resolved models plus Claude Code cost estimates.
Live Reviews Use Metered Calls
The live run invokes Claude Code seven times. Every review call has a client-side cap of $0.25.
The recorded total_cost_usd values are client-side estimates. They can differ from the final bill.
Expect several minutes of quiet while the seven reviewers run sequentially.
Before you run the live evaluation, predict whether the requested aliases plus resolved model names will always match.
- Run the full evaluation with live reviews by running:
python evaluate.py --reviews
What Should the Live Evaluation Show?
- The terminal should print all seven cases followed by the updated summary.
- The live_reviews_enabled value should show true.
- The requested-model usage should count the Haiku plus Sonnet plus Opus aliases selected by the reviewer profiles.
- The resolved-model usage should list the actual model names reported by Claude Code.
- The Claude Code estimated cost should aggregate the value returned by every completed review.
- Return to reports/evaluation.json in the Explorer sidebar.
- Confirm that the summary object contains routing quality plus reviewer usage plus model usage plus both cost totals.
- Confirm that each case includes its decision plus its live review result.
Live Evaluation Stops Early?
Confirm that the current terminal still has the active virtual environment plus the TypeSafe API key. Check the authenticated Claude Code status from the setup step if reviewer dispatch fails.
Help me debug the live evaluation.
That completes the core pipeline. You now have a labeled report that exposes routing mistakes plus reviewer allocation plus resolved model usage plus estimated cost.
Secret mission
Defeat the Three-Line Authorization Trap
A three-line authorization regression can carry more risk than a much larger routine change. Build an adversarial policy check that proves misleading low-risk signals cannot bypass the high-risk reviewer.
Clean Up Your Resources
Clean Up Your Resources
Choose how much of your local review pipeline to keep. Only new Jev requests or Claude Code reviews can create metered usage.
Resources you used:
- Local project folder. The jev-review-router folder contains your source files. It contains your project-scoped Claude Code reviewer profiles. It includes the .venv Python virtual environment. It also includes your JSON evaluation reports.
- Current shell credential. The TypeSafe API key is exported in your current shell as TYPESAFE_API_KEY.
Keep everything running
No action is needed for the project files. Choose this if you plan to keep testing the router or extending its safety policy.
- Retain the jev-review-router folder with its source files.
- Keep reports/evaluation.json as evidence of your routing results.
- Close the current terminal when you finish working.
- Export TYPESAFE_API_KEY again only when you want to run a new Jev request.
Pause - I'll come back to this later
Pause the shell environment while keeping every project file. Your router stays ready for the next session.
- Clear the TypeSafe API key from the current shell by running this command:
unset TYPESAFE_API_KEY
What Does This Command Do?
This removes TYPESAFE_API_KEY from the current shell. It does not delete the key stored in your TypeSafe account.
- Leave the virtual environment by running this command:
deactivate
Why Deactivate First?
The deactivate command returns your shell to its normal Python environment. Your .venv files remain available for the next session.
- Omit --reviews from your next evaluation if you want to avoid Claude Code reviewer calls.
Your local pipeline is now paused safely. No new metered calls occur while the evaluation commands stay idle.
Delete - I don't want to use this again
Remove every locally teardownable project resource. The remote TypeSafe key sits outside the verified cleanup path.
- Clear the TypeSafe API key from the current shell by running this command:
unset TYPESAFE_API_KEY
What About the TypeSafe Key?
This command removes the credential only from your current shell. TypeSafe's public documentation does not specify a verified revoke or delete workflow for the account key.
- Leave the virtual environment by running this command:
deactivate
Why Leave the Environment?
This returns your shell to its normal Python environment before the project folder disappears. The disposable .venv remains inside the folder until deletion.
Your shell is now clear of the project environment. The next action permanently removes the local pipeline.
Permanent Local Deletion
Deleting the folder is permanent.
- Pause here if you may want the code or evaluation report later.
- Return to the project terminal from earlier.
- Move out of the project folder and delete it by running these commands:
cd ..
rm -rf jev-review-router
What Do These Commands Remove?
The first line moves your shell to the folder containing the project. The second line permanently removes the complete local project folder.
- Search your Mac for jev-review-router to verify the deletion.
That closes the local project cleanly. Your search should return no jev-review-router folder.
Nice Work!
Nice Work!
You made it! Your adaptive code review router now turns typed Jev evidence into deterministic routes for progressive Claude Code reviewers.
What you learned:
- Built a typed Jev evidence layer with one atomic question for implementation complexity. Added another atomic question for operational or security consequence. Preserved confidence for policy decisions. Preserved probabilities for inspection.
- Owned final routing through deterministic policy written in Python. Applied confidence escalation. Added non-negotiable safety gates. Matched routine reviews to Haiku. Matched complex reviews to Sonnet. Matched high-risk reviews to Opus. Kept every Claude Code reviewer read-only.
- Produced a labeled evaluation report for routing accuracy. Counted over-escalations. Counted under-escalations. Recorded reviewer usage. Recorded resolved model usage. Tracked Jev cost. Tracked Claude Code estimated cost.
- Secret Mission: Built a fail-closed authorization gate. Proved that a three-line authorization regression still reaches high_risk under misleading low-risk signals. Confirmed the invariant with verify_safety_gate.py.
Ready to quiz yourself?