Build a Safe Jev Incident Router

Build a fail-safe Jev incident router with deterministic Python policy.

Introduction

30 Second Summary

When a service breaks, the first routing decision determines who investigates the problem. A confident wrong answer can waste precious time while customers keep feeling the outage.

In this project, you will build a command-line incident router that asks Jev for a specialist recommendation through OpenRouter using the official TypeSafe Python SDK. Deterministic Python safety rules approve or reject that recommendation before a read-only Claude Code investigator acts.

What You'll Build

Run an incident from macOS Terminal to watch an approved route name a read-only specialist while every uncertain or broken outcome stops at human review.

By the end of this project, you'll have:

  • Run a live incident through OpenRouter. Inspect the selected route plus confidence. Review the serving model plus token usage. See reported cost when available.
  • Replay a recorded response without a network request. Use pytest to get the same policy result every time.
  • Launch an approved read-only investigator for application, deployment, or platform evidence. Keep its access limited to file inspection in plan mode.
  • Secret Mission: uncover a hidden authentication outage that looks safe to Jev. Prove Python still sends it to human review.

Are there any prerequisites?

This project assumes an existing OpenRouter API key with available credit. You also need an authenticated Claude Code installation. Your macOS environment needs Python 3.10 or newer.

Before We Start

You are committing to a router where Jev provides a probabilistic recommendation while Python makes the final fail-closed decision. Its Claude Code investigators remain read-only so evidence gathering cannot change the project.

Set Up the Pinned Python Workspace

The incident router depends on stable Python library behavior. A mismatched environment can break imports before your safety policy even runs.

This step creates a pinned workspace for the TypeSafe Python SDK plus pytest.

You also verify access to OpenRouter. Finally, you confirm that Claude Code is ready for the read-only investigation step.

In this step, get ready to:
  • Verify a compatible Python installation on macOS.
  • Create an isolated workspace with pinned dependencies.
  • Confirm your existing OpenRouter key and Claude Code login.
Verify Python on macOS

The TypeSafe Python SDK and pytest require Python 3.10 or newer. Checking the runtime first prevents confusing installation failures later.

  • Press Cmd+Space to open Spotlight.
  • Type Terminal into the search field.
  • Press Enter to open Terminal.
  • Check your installed Python version by running this command:
python3 --version

What does this command check?

The command asks the active python3 executable to report its version. The result determines whether you can use the pinned packages.

✔️ I see version 3.10 or higher

Your Python runtime meets the project requirement. Keep this Terminal window open for the workspace setup.

ⓧ I see an older version

Your current Python runtime is below the package requirement. Install Python 3.14.8 before creating the virtual environment.

  • Visit the official Python downloads page for macOS.
  • Download the Python 3.14.8 macOS installer.
  • Complete the installation using the installer prompts.
  • Quit Terminal after the installation finishes.
  • Press Cmd+Space to reopen Spotlight.
  • Type Terminal into the search field.
  • Press Enter to reopen Terminal.
  • Repeat the version check from above. Confirm that it now reports Python 3.10 or newer.

ⓧ Command not found

Your Mac cannot currently find a Python installation. Install Python 3.14.8 from the official source.

  • Visit the official Python downloads page for macOS.
  • Download the Python 3.14.8 macOS installer.
  • Complete the installation using the installer prompts.
  • Quit Terminal after the installation finishes.
  • Press Cmd+Space to reopen Spotlight.
  • Type Terminal into the search field.
  • Press Enter to reopen Terminal.
  • Repeat the version check from above. Confirm that it now reports Python 3.10 or newer.
Create the isolated workspace

A virtual environment keeps this router's packages separate from the rest of your Mac. Exact dependency pins make later fixture replays easier to reproduce.

  • Move to your Desktop by running this command:
cd ~/Desktop

Why start on the Desktop?

The Desktop gives the workspace a predictable location. You can find the project without searching through your home directory.

  • Create the fail-safe-incident-router folder by running these commands:
mkdir fail-safe-incident-router
cd fail-safe-incident-router

What do these commands do?

  • The first command creates the fail-safe-incident-router folder on your Desktop.
  • The second command makes that folder the active Terminal location.

Does the folder already exist?

A folder-exists message means the workspace is already on your Desktop. Run the second command again to enter it.

Ask for help if Terminal still points somewhere unexpected: help me enter the fail-safe-incident-router folder on my Desktop

The requirements.txt file records the exact packages that belong in the environment. This prevents a future install from silently selecting different releases.

  • Create requirements.txt with the built-in nano editor by running this command:
nano requirements.txt

What does nano do?

Nano opens a plain-text editor inside Terminal. Supplying a new filename creates that file when you save it.

  • Enter the exact dependency pins below:
typesafe-sdk==0.7.2
pytest==9.1.1

Why pin both versions?

  • The typesafe-sdk==0.7.2 pin keeps the router on the SDK interface used by this project.
  • The pytest==9.1.1 pin keeps the later evaluation suite reproducible.
  • Save requirements.txt by pressing Ctrl+O.
  • Accept the filename by pressing Enter.

Nano reports that it wrote two lines. Your dependency pins are now stored on disk.

  • Exit nano by pressing Ctrl+X.
  • Confirm the saved dependency pins by running this command:
cat requirements.txt

What should this print?

The command displays the saved file without changing it. You should see the two exact package pins from the code block above.

The .gitignore file keeps generated environment files out of future version control. It also excludes Python cache files created during later runs.

  • Create .gitignore with nano by running this command:
nano .gitignore

Why does the filename start with a dot?

A leading dot makes this a hidden configuration file on macOS. Git reads it to identify generated files that should remain outside commits.

  • Enter the exact ignore rules below:
.venv/
__pycache__/
.pytest_cache/
*.pyc

What do these rules exclude?

  • The .venv/ rule excludes installed project dependencies.
  • The remaining rules exclude Python bytecode plus pytest cache data.
  • Save .gitignore by pressing Ctrl+O.
  • Accept the filename by pressing Enter.

Nano reports that it wrote four lines. Your generated Python files now have explicit ignore rules.

  • Exit nano by pressing Ctrl+X.
  • Confirm the saved ignore rules by running this command:
cat .gitignore

What should this print?

The command prints the four ignore rules from the code block above. This confirms that the hidden file exists with the correct contents.

  • Create the virtual environment by running these commands:
python3 -m venv .venv
source .venv/bin/activate

How does the environment stay isolated?

  • The first command creates a local Python environment inside .venv.
  • The second command directs this Terminal session to use that environment.

Your Terminal prompt now starts with (.venv). That prefix confirms the isolated environment is active.

Missing the environment prefix?

Confirm that Terminal is still inside fail-safe-incident-router. Run the activation command again from that folder.

Ask for help if activation still fails: help me activate the .venv environment in my fail-safe-incident-router folder

The environment is empty until pip reads the version pins. Installation now places both packages inside .venv.

  • Install the pinned dependencies by running this command:
python3 -m pip install -r requirements.txt

What does this installation use?

Pip reads each exact version from requirements.txt. Because the virtual environment is active, the packages stay inside this project.

The installation may take a minute while pip downloads the packages. When it finishes, you return to the (.venv) prompt.

  • Inspect both installed packages by running this command:
python3 -m pip show typesafe-sdk pytest

What does the package report prove?

Each package section includes its installed version plus its location. The location should point inside the project's .venv folder.

You should see TypeSafe SDK version 0.7.2 plus pytest version 9.1.1.

Did the installation fail?

Check that your prompt begins with (.venv). A missing prefix means the packages may be installing outside the project environment.

Ask for help with the exact terminal output: help me install the pinned TypeSafe SDK and pytest packages inside my active .venv

Why use TypeSafe with OpenRouter?

The official TypeSafe Python SDK preserves the typed Choice plus Noul decision model used by the router. OpenRouter lets the project use your existing credential instead of adding another account.

✔️ Awesome, I've got everything!

Great. Double check that both project files are saved inside fail-safe-incident-router.

ⓧ I'd like to double check the full code

Compare your two setup files with these exact versions.

typesafe-sdk==0.7.2
pytest==9.1.1
.venv/
__pycache__/
.pytest_cache/
*.pyc
Confirm credentials and tools

The router needs the OpenRouter key in its environment when it makes a live request. A presence check can confirm that setup without exposing the credential.

  • Check for the existing key without printing its value by running this command:
if [[ -n $OPENROUTER_API_KEY ]]
then
  echo "OPENROUTER_API_KEY is set"
else
  echo "OPENROUTER_API_KEY is missing"
fi

How does this protect the key?

The conditional checks whether OPENROUTER_API_KEY contains a value. It prints only the presence status.

You should see OPENROUTER_API_KEY is set. Your credential remains hidden.

Is the key missing?

Reload the shell configuration used for your existing OpenRouter setup. Repeat the safe presence check after the key is available in this Terminal session.

Ask for guidance without sharing the credential: help me make my existing OpenRouter key available in this macOS Terminal session without printing it

Package metadata proves that files were installed. An import check proves that Python can actually load the SDK objects used by the router.

  • Confirm the pinned imports by running this command:
python3 -c "from typesafe_sdk import Choice, Noul, TypeSafeClient
import pytest
print('Pinned imports are ready')"

What does the import check cover?

  • The first import loads the three TypeSafe SDK objects used by the router.
  • The second import confirms that pytest is available for the evaluation suite.

You should see Pinned imports are ready. That message only prints after every import succeeds.

Seeing an import error?

Check that your prompt starts with (.venv). Repeat the dependency installation inside the active environment if the prefix was missing.

Ask for help with the module named by the error: help me fix the pinned import check inside my active Python virtual environment

Claude Code must be current enough to support the project-scoped read-only agents used later. Its authentication check confirms that the existing installation can start an authenticated session.

  • Check the installed Claude Code version plus its authentication status by running these commands:
claude -v
claude auth status

What do these checks confirm?

  • The first command reports the installed Claude Code version.
  • The second command reports whether Claude Code has an active login.

✔️ I see version 2.1.293 or newer

Claude Code is version-compatible. A logged-in authentication status confirms that the existing setup is ready.

ⓧ I see an older version

Update Claude Code before creating project-scoped investigators. The update keeps the later agent configuration aligned with the documented CLI behavior.

  • Update Claude Code by running this command:
claude update

What does this command change?

The command updates the existing Claude Code installation to the latest available release. It keeps your current project files untouched.

  • Repeat the version check from above after the update finishes.
  • Confirm that the reported version is 2.1.293 or newer.

ⓧ I am not logged in

Your installation is available without an active authenticated session. The login flow opens a browser so you can reconnect the existing account.

  • Start the Claude Code login flow by running this command:
claude auth login

What happens during login?

Claude Code opens its browser authentication flow. Complete the prompts using the account you already use for Claude Code.

  • Complete the browser authentication prompts.
  • Repeat the authentication check from above.
  • Confirm that the result shows a logged-in status.

Before you run the final checks, which versions do you expect the active environment to report?

  • Verify the complete workspace setup by running these commands:
python3 --version
python3 -m pip show typesafe-sdk pytest
claude -v
claude auth status

What does the final check cover?

These commands verify the Python runtime plus both pinned packages. They also verify the Claude Code version plus its authentication state.

You should see Python 3.10 or newer. The package entries should show TypeSafe SDK 0.7.2 plus pytest 9.1.1.

Claude Code should report version 2.1.293 or newer. Its authentication output should show that you are logged in.

That closes the setup risk. Your pinned workspace can now make its first live Jev decision through OpenRouter.

Route a Live Incident with Jev

Your pinned Python workspace can now reach the services this router depends on. The next goal is a visible decision from Jev through OpenRouter.

This step uses the official TypeSafe Python SDK to ask typed routing questions. You will record the raw response before testing how the first routing draft handles uncertain output.

In this step, get ready to:
  • Build the initial typed incident router.
  • Record a live Jev response for a cache latency incident.
  • Replay a low-confidence response to expose the routing gap.
Build the typed router

A System One request accepts structured state plus typed questions. The router uses Choice to select an investigator route.

The Noul question returns the probability that human review is needed. This initial draft prints both answers as audit metadata.

  • Create incident_router.py inside the open project workspace using your code editor.
  • Use the complete file in the double-check tab as the initial router implementation.

✔️ I created the router file

Keep incident_router.py saved in the open project workspace.

ⓧ I'd like to double check the full code

Compare your complete incident_router.py file with this initial router.

from __future__ import annotations

import argparse
import json
import os
from dataclasses import asdict, dataclass
from pathlib import Path
from typing import Any, Callable

from typesafe_sdk import Choice, Noul, TypeSafeClient

MODEL = "jev-1.13"
OPENROUTER_BASE_URL = "https://openrouter.ai/api"

AGENT_BY_ROUTE = {
    "application": "application-investigator",
    "deployment": "deployment-investigator",
    "platform": "platform-investigator",
}

QUESTIONS = {
    "investigator": Choice(
        instructions="Which read-only investigator should examine this incident?",
        criteria={
            "application": "Application behavior, request handling, or business logic.",
            "deployment": "A release, configuration change, or rollout problem.",
            "platform": "Infrastructure, networking, capacity, or shared platform behavior.",
        },
    ),
    "human_review": Noul(
        instructions=(
            "Does this incident require immediate human review because it may involve "
            "authentication, security, data risk, broad impact, or ambiguous evidence?"
        ),
        criteria={
            "true": "Human judgment is required before automated investigation.",
            "false": "A read-only specialist can safely investigate first.",
        },
    ),
}

Payload = dict[str, Any]
PayloadProvider = Callable[[dict[str, Any]], Payload]


@dataclass(frozen=True)
class JevDecision:
    route: str
    route_confidence: float
    human_review_probability: float
    model: str
    input_tokens: int | None
    output_tokens: int | None
    cost: float | None


@dataclass(frozen=True)
class RoutingResult:
    action: str
    reason: str
    agent: str | None
    model: str | None
    route_confidence: float | None
    human_review_probability: float | None
    input_tokens: int | None
    output_tokens: int | None
    cost: float | None


def load_json(path: Path) -> Payload:
    value = json.loads(path.read_text(encoding="utf-8"))
    if not isinstance(value, dict):
        raise ValueError("JSON root must be an object")
    return value


def save_json(path: Path, payload: Payload) -> None:
    path.parent.mkdir(parents=True, exist_ok=True)
    path.write_text(json.dumps(payload, indent=2) + "\n", encoding="utf-8")


def call_jev(incident: dict[str, Any]) -> Payload:
    api_key = os.environ["OPENROUTER_API_KEY"]
    with TypeSafeClient(
        api_key=api_key,
        base_url=OPENROUTER_BASE_URL,
        model=MODEL,
    ) as client:
        response = client.system_one(state=incident, questions=QUESTIONS)
        raw = response.raw_http_response.json()
    if not isinstance(raw, dict):
        raise ValueError("Jev response root must be an object")
    return raw


def decision_from_payload(payload: Payload) -> JevDecision:
    model = payload.get("model")
    answers = payload.get("answers")
    route_answer = answers.get("investigator")
    review_answer = answers.get("human_review")

    route = route_answer.get("choice")
    confidence = route_answer.get("confidence")
    review_probability = review_answer.get("noul")
    usage = payload.get("usage")

    return JevDecision(
        route=route,
        route_confidence=float(confidence),
        human_review_probability=float(review_probability),
        model=model,
        input_tokens=usage.get("input_tokens"),
        output_tokens=usage.get("output_tokens"),
        cost=usage.get("cost"),
    )


def enforce_policy(
    incident: dict[str, Any], decision: JevDecision
) -> RoutingResult:
    return RoutingResult(
        action="dispatch_read_only_investigator",
        reason="policy_approved",
        agent=AGENT_BY_ROUTE[decision.route],
        model=decision.model,
        route_confidence=decision.route_confidence,
        human_review_probability=decision.human_review_probability,
        input_tokens=decision.input_tokens,
        output_tokens=decision.output_tokens,
        cost=decision.cost,
    )


def route_incident(
    incident: dict[str, Any], provider: PayloadProvider = call_jev
) -> RoutingResult:
    payload = provider(incident)
    decision = decision_from_payload(payload)
    return enforce_policy(incident, decision)


def build_parser() -> argparse.ArgumentParser:
    parser = argparse.ArgumentParser(description="Fail-safe Jev incident router")
    parser.add_argument("incident", type=Path)
    parser.add_argument("--response-fixture", type=Path)
    parser.add_argument("--record-response", type=Path)
    return parser


def main() -> None:
    args = build_parser().parse_args()
    if args.response_fixture and args.record_response:
        raise SystemExit("Choose either --response-fixture or --record-response")

    incident = load_json(args.incident)

    if args.response_fixture:
        recorded_payload = load_json(args.response_fixture)
        provider: PayloadProvider = lambda _incident: recorded_payload
    elif args.record_response:
        def recording_provider(state: dict[str, Any]) -> Payload:
            payload = call_jev(state)
            save_json(args.record_response, payload)
            return payload
        provider = recording_provider
    else:
        provider = call_jev

    result = route_incident(incident, provider)
    print(json.dumps(asdict(result), indent=2))


if __name__ == "__main__":
    main()

How does the router work?

  • The QUESTIONS mapping gives Jev one investigator choice plus one human-review probability question.
  • The call_jev() function sends the incident state through the configured TypeSafe client.
  • The --record-response option saves the raw response before the router prints its audit result.
  • The --response-fixture option replaces the live provider with a recorded payload.
  • Save incident_router.py.
  • Confirm the router imports with the pinned model by running:
python3 -c "import incident_router; print(incident_router.MODEL)"

What does this check prove?

Python imports the complete module before printing jev-1.13. A successful result proves the file has valid syntax plus accessible SDK imports.

You should see jev-1.13 in the terminal.

Router import failing?

Confirm that the terminal prompt still shows the active .venv environment. An inactive environment can hide the installed TypeSafe package.

Compare the import line with from typesafe_sdk import Choice, Noul, TypeSafeClient.

Help me fix the incident router import error.

Record a live decision

The router needs a concrete incident state before Jev can choose a specialist. This cache latency example contains impact details plus recent operational evidence.

  • Create the incidents directory inside the open project workspace.
  • Create cache_latency.json inside the incidents directory with this content:
{
  "incident_id": "INC-1042",
  "summary": "Checkout requests became slow after a cache configuration rollout",
  "customer_impact": "Elevated latency for a subset of checkout requests",
  "recent_change": "Cache time-to-live configuration changed 20 minutes ago",
  "evidence": [
    "Application error rate is unchanged",
    "Cache miss rate increased",
    "Database CPU remains within its normal range"
  ],
  "policy_flags": {
    "authentication_outage": false
  }
}

What does this incident describe?

The incident ties higher checkout latency to a recent cache configuration rollout. Its evidence gives Jev signals for comparing application, deployment, and platform routes.

The authentication_outage flag is false for this incident. The deterministic policy uses that signal in the next step.

  • Save incidents/cache_latency.json.

This next command makes one paid model request through your existing OpenRouter billing arrangement. It records the raw response without printing your API key.

Before you run it, form a prediction about which investigator route best fits the cache evidence.

  • Route the live incident while recording its raw response by running:
python3 incident_router.py incidents/cache_latency.json --record-response tests/fixtures/live_cache_latency.json

What does this command do?

  • The router loads incidents/cache_latency.json as the request state.
  • The TypeSafe client sends the typed questions to the pinned model through OpenRouter.
  • The recording provider writes the untouched response to tests/fixtures/live_cache_latency.json.
  • The router prints the selected agent plus its audit metadata.

You should see JSON containing action, agent, model, route_confidence, human_review_probability, token usage, and cost when OpenRouter reports it.

That is your first live incident decision. The saved response now gives you a replayable record of the exact model output.

Live request failing?

Confirm that OPENROUTER_API_KEY remains present in this terminal without printing its value. A new terminal session may not inherit the earlier environment.

Check that OPENROUTER_BASE_URL is exactly https://openrouter.ai/api.

Help me diagnose the live Jev request safely.

Replay a low-confidence recommendation

A recorded fixture makes one model outcome repeatable. This fixture is syntactically valid while assigning only 0.61 confidence to the deployment route.

  • Create the tests/fixtures directories inside the open project workspace.
  • Create low_confidence.json inside tests/fixtures with this recorded test double:
{
  "fixture_source": "authored_test_double",
  "model": "typesafe/jev-1.13-20260917",
  "answers": {
    "investigator": {
      "type": "choice",
      "choice": "deployment",
      "confidence": 0.61,
      "probabilities": {
        "application": 0.18,
        "deployment": 0.61,
        "platform": 0.21
      }
    },
    "human_review": {
      "type": "noul",
      "noul": 0.15
    }
  },
  "usage": {
    "input_tokens": 302,
    "output_tokens": 24,
    "cost": 0.000013
  }
}

Why use a recorded fixture?

The fixture preserves a specific Jev-shaped response without making another network request. Replaying the same payload isolates the router's Python behavior.

The fixture_source field identifies this payload as an authored test double. It remains separate from the live response you recorded.

  • Save tests/fixtures/low_confidence.json.

Before you replay it, predict whether this initial router accepts or escalates the 0.61 recommendation.

  • Replay the low-confidence response through the initial router by running:
python3 incident_router.py incidents/cache_latency.json --response-fixture tests/fixtures/low_confidence.json

What does this replay test?

The --response-fixture option bypasses the live provider. The router receives the same deployment recommendation on every run.

This keeps model variation out of the check. Any resulting action comes from the current Python routing flow.

You will see dispatch_read_only_investigator with deployment-investigator even though route_confidence is only 0.61.

This shortfall is intentional

The first router converts a valid model response directly into a dispatch. It has no deterministic confidence threshold controlling the final action.

The replay proves that syntactic validity alone cannot make a recommendation safe. Python needs explicit authority over the outcome.

Fixture replay failing?

Confirm that the file path is exactly tests/fixtures/low_confidence.json. The router resolves that path from the open project workspace.

Compare the fixture punctuation with the supplied JSON. A missing comma or brace prevents load_json() from reading it.

Help me fix the recorded fixture replay.

Your live router can now call Jev plus replay recorded evidence. Next, you will make deterministic Python policy overrule risky or uncertain recommendations.

Make Python the Fail-Safe Authority

The previous step gave you a live Jev recommendation. Your replay also showed that a valid response with 0.61 confidence could still select an investigator.

Python now becomes the final authority. It validates every response before deterministic policy decides whether an investigator is safe to dispatch.

In this step, get ready to:
  • Validate the model identity plus every decision field.
  • Escalate unavailable or unsafe outcomes to human review.
  • Replay low-confidence and invalid fixtures without a network request.
Validate every Jev response

A syntactically valid JSON object can still contain impossible probabilities or an unexpected model. Strict parsing converts only trusted response shapes into a typed decision.

What does strict parsing protect?

The parser accepts serving models that begin with typesafe/jev-1.13. This allows the dated serving suffix while preserving the pinned model family.

It also checks the route label plus every probability. Token counts must be nonnegative integers when present.

  • Switch back to incident_router.py in your editor.
  • Select the complete file using Cmd+A on macOS or Ctrl+A on Windows.
  • Open the second tab below to reveal the completed fail-safe router.
  • Replace the selected content with the exact full file from that tab.
  • Save incident_router.py.

✔️ I've updated the router

Your router now separates the model recommendation from the policy decision. Continue to the invalid-response fixture.

ⓧ I'd like to double check the full code

Compare your complete incident_router.py file with this version. Every identifier must match exactly.

from __future__ import annotations

import argparse
import json
import math
import os
from dataclasses import asdict, dataclass
from pathlib import Path
from typing import Any, Callable

from typesafe_sdk import Choice, Noul, TypeSafeClient

MODEL = "jev-1.13"
OPENROUTER_BASE_URL = "https://openrouter.ai/api"
MIN_ROUTE_CONFIDENCE = 0.80
HUMAN_REVIEW_THRESHOLD = 0.50

AGENT_BY_ROUTE = {
    "application": "application-investigator",
    "deployment": "deployment-investigator",
    "platform": "platform-investigator",
}

QUESTIONS = {
    "investigator": Choice(
        instructions="Which read-only investigator should examine this incident?",
        criteria={
            "application": "Application behavior, request handling, or business logic.",
            "deployment": "A release, configuration change, or rollout problem.",
            "platform": "Infrastructure, networking, capacity, or shared platform behavior.",
        },
    ),
    "human_review": Noul(
        instructions=(
            "Does this incident require immediate human review because it may involve "
            "authentication, security, data risk, broad impact, or ambiguous evidence?"
        ),
        criteria={
            "true": "Human judgment is required before automated investigation.",
            "false": "A read-only specialist can safely investigate first.",
        },
    ),
}

Payload = dict[str, Any]
PayloadProvider = Callable[[dict[str, Any]], Payload]


@dataclass(frozen=True)
class JevDecision:
    route: str
    route_confidence: float
    human_review_probability: float
    model: str
    input_tokens: int | None
    output_tokens: int | None
    cost: float | None


@dataclass(frozen=True)
class RoutingResult:
    action: str
    reason: str
    agent: str | None
    model: str | None
    route_confidence: float | None
    human_review_probability: float | None
    input_tokens: int | None
    output_tokens: int | None
    cost: float | None


def load_json(path: Path) -> Payload:
    value = json.loads(path.read_text(encoding="utf-8"))
    if not isinstance(value, dict):
        raise ValueError("JSON root must be an object")
    return value


def save_json(path: Path, payload: Payload) -> None:
    path.parent.mkdir(parents=True, exist_ok=True)
    path.write_text(json.dumps(payload, indent=2) + "\n", encoding="utf-8")


def call_jev(incident: dict[str, Any]) -> Payload:
    api_key = os.environ["OPENROUTER_API_KEY"]
    with TypeSafeClient(
        api_key=api_key,
        base_url=OPENROUTER_BASE_URL,
        model=MODEL,
    ) as client:
        response = client.system_one(state=incident, questions=QUESTIONS)
        raw = response.raw_http_response.json()
    if not isinstance(raw, dict):
        raise ValueError("Jev response root must be an object")
    return raw


def is_probability(value: object) -> bool:
    return (
        isinstance(value, (int, float))
        and not isinstance(value, bool)
        and math.isfinite(float(value))
        and 0.0 <= float(value) <= 1.0
    )


def optional_nonnegative_int(value: object, field: str) -> int | None:
    if value is None:
        return None
    if isinstance(value, bool) or not isinstance(value, int) or value < 0:
        raise ValueError(f"{field} must be a nonnegative integer or null")
    return value


def optional_nonnegative_number(value: object, field: str) -> float | None:
    if value is None:
        return None
    if (
        isinstance(value, bool)
        or not isinstance(value, (int, float))
        or not math.isfinite(float(value))
        or float(value) < 0.0
    ):
        raise ValueError(f"{field} must be a nonnegative number or null")
    return float(value)


def decision_from_payload(payload: Payload) -> JevDecision:
    model = payload.get("model")
    if not isinstance(model, str) or not model.startswith("typesafe/jev-1.13"):
        raise ValueError("Unexpected serving model")

    answers = payload.get("answers")
    if not isinstance(answers, dict):
        raise ValueError("answers must be an object")

    route_answer = answers.get("investigator")
    review_answer = answers.get("human_review")
    if not isinstance(route_answer, dict) or route_answer.get("type") != "choice":
        raise ValueError("Missing choice answer")
    if not isinstance(review_answer, dict) or review_answer.get("type") != "noul":
        raise ValueError("Missing noul answer")

    route = route_answer.get("choice")
    confidence = route_answer.get("confidence")
    probabilities = route_answer.get("probabilities")
    review_probability = review_answer.get("noul")

    if route not in AGENT_BY_ROUTE:
        raise ValueError("Unknown route")
    if not is_probability(confidence) or not is_probability(review_probability):
        raise ValueError("Invalid confidence or review probability")
    if not isinstance(probabilities, dict) or set(probabilities) != set(AGENT_BY_ROUTE):
        raise ValueError("Invalid route probabilities")
    if not all(is_probability(value) for value in probabilities.values()):
        raise ValueError("Invalid route probability value")

    usage = payload.get("usage")
    if not isinstance(usage, dict):
        raise ValueError("usage must be an object")

    return JevDecision(
        route=route,
        route_confidence=float(confidence),
        human_review_probability=float(review_probability),
        model=model,
        input_tokens=optional_nonnegative_int(usage.get("input_tokens"), "input_tokens"),
        output_tokens=optional_nonnegative_int(usage.get("output_tokens"), "output_tokens"),
        cost=optional_nonnegative_number(usage.get("cost"), "cost"),
    )


def human_review(reason: str, decision: JevDecision | None = None) -> RoutingResult:
    return RoutingResult(
        action="human_review",
        reason=reason,
        agent=None,
        model=decision.model if decision else None,
        route_confidence=decision.route_confidence if decision else None,
        human_review_probability=(
            decision.human_review_probability if decision else None
        ),
        input_tokens=decision.input_tokens if decision else None,
        output_tokens=decision.output_tokens if decision else None,
        cost=decision.cost if decision else None,
    )


def enforce_policy(
    incident: dict[str, Any], decision: JevDecision
) -> RoutingResult:
    policy_flags = incident.get("policy_flags")
    if (
        isinstance(policy_flags, dict)
        and policy_flags.get("authentication_outage") is True
    ):
        return human_review("authentication_outage_policy", decision)

    if decision.human_review_probability >= HUMAN_REVIEW_THRESHOLD:
        return human_review("high_human_review_probability", decision)

    if decision.route_confidence < MIN_ROUTE_CONFIDENCE:
        return human_review("low_route_confidence", decision)

    return RoutingResult(
        action="dispatch_read_only_investigator",
        reason="policy_approved",
        agent=AGENT_BY_ROUTE[decision.route],
        model=decision.model,
        route_confidence=decision.route_confidence,
        human_review_probability=decision.human_review_probability,
        input_tokens=decision.input_tokens,
        output_tokens=decision.output_tokens,
        cost=decision.cost,
    )


def route_incident(
    incident: dict[str, Any], provider: PayloadProvider = call_jev
) -> RoutingResult:
    try:
        payload = provider(incident)
    except Exception:
        return human_review("jev_unavailable")

    try:
        decision = decision_from_payload(payload)
    except (KeyError, TypeError, ValueError):
        return human_review("invalid_jev_response")

    return enforce_policy(incident, decision)


def build_parser() -> argparse.ArgumentParser:
    parser = argparse.ArgumentParser(description="Fail-safe Jev incident router")
    parser.add_argument("incident", type=Path)
    parser.add_argument("--response-fixture", type=Path)
    parser.add_argument("--record-response", type=Path)
    return parser


def main() -> None:
    args = build_parser().parse_args()
    if args.response_fixture and args.record_response:
        raise SystemExit("Choose either --response-fixture or --record-response")

    incident = load_json(args.incident)

    if args.response_fixture:
        recorded_payload = load_json(args.response_fixture)
        provider: PayloadProvider = lambda _incident: recorded_payload
    elif args.record_response:
        def recording_provider(state: dict[str, Any]) -> Payload:
            payload = call_jev(state)
            save_json(args.record_response, payload)
            return payload
        provider = recording_provider
    else:
        provider = call_jev

    result = route_incident(incident, provider)
    print(json.dumps(asdict(result), indent=2))


if __name__ == "__main__":
    main()

How does the completed router work?

  • The frozen JevDecision stores only validated model output.
  • The frozen RoutingResult keeps the final action plus its audit metadata.
  • The route_incident() function converts provider failures into jev_unavailable.
  • The enforce_policy() function applies Python thresholds before selecting an agent.
Add an invalid response fixture

A response fixture lets you test malformed model output repeatedly. This fixture uses an impossible confidence of 1.4 to prove that invalid data fails closed.

  • Expand the tests/fixtures folder in your editor's file sidebar.
  • Click the new-file control in the sidebar.
  • Name the file invalid.json.
  • Paste this recorded Jev-shaped response into invalid.json:
{
  "fixture_source": "authored_test_double",
  "model": "typesafe/jev-1.13-20260917",
  "answers": {
    "investigator": {
      "type": "choice",
      "choice": "platform",
      "confidence": 1.4,
      "probabilities": {
        "application": 0.1,
        "deployment": 0.1,
        "platform": 0.8
      }
    },
    "human_review": {
      "type": "noul",
      "noul": 0.1
    }
  },
  "usage": {
    "input_tokens": 300,
    "output_tokens": 24,
    "cost": 0.000013
  }
}

Why is this response invalid?

The route label belongs to AGENT_BY_ROUTE. The answer objects also use the expected types.

The confidence falls outside the accepted range from 0.0 to 1.0. That single invalid field rejects the complete decision.

  • Save tests/fixtures/invalid.json.
Replay fail-closed outcomes

The recorded-response provider is a form of dependency injection. It supplies fixed input to the same routing policy without contacting the live service.

Before you replay the low-confidence fixture, do you think its valid deployment recommendation still reaches an investigator?

  • Return to the activated terminal from earlier.
  • Replay the low-confidence response by running:
python3 incident_router.py incidents/cache_latency.json --response-fixture tests/fixtures/low_confidence.json

What should you see?

You see human_review as the action. The reason is low_route_confidence.

The route confidence remains visible as 0.61 for auditing. The mapped investigator is withheld.

That is the policy boundary working. Jev's recommendation remains useful evidence while Python controls the action.

Before you replay the invalid fixture, do you expect the router to preserve its model metadata or reject the decision completely?

  • Replay the invalid response by running:
python3 incident_router.py incidents/cache_latency.json --response-fixture tests/fixtures/invalid.json

How does invalid data fail closed?

You see human_review with the reason invalid_jev_response. The decision metadata is null because the payload never became a trusted JevDecision.

Fixture replay not working?

Confirm that the activated terminal is still inside the workspace containing incident_router.py. Check that both fixture paths match the files under tests/fixtures.

Compare incident_router.py with the complete version above if the old deployment action still appears.

Help me diagnose my fixture replay.

Your router now fails closed when the model is uncertain or malformed. Next, you will give approved routes a set of read-only Claude Code investigators.

Create Read-Only Claude Investigators

Your Python policy now blocks unavailable, invalid, risky, or uncertain decisions from reaching an investigator. Approved routes can move forward with their audit metadata intact.

A safe route still needs a safe investigator. In this step, you will give Claude Code three project-scoped roles that can inspect evidence without modifying files or running shell commands.

In this step, get ready to:
  • Create application, deployment, and platform investigator definitions.
  • Restrict every investigator to read-only file tools in plan mode.
  • Run the platform investigator against the cache-latency incident.
Define the application scope

A project-scoped agent stores its role inside .claude/agents/. Its frontmatter controls the available tools before the investigation begins.

Why project-scoped agents?

Project-scoped agents keep the investigation rules beside the router they protect. Anyone who uses this workspace receives the same evidence requirements.

The tool allowlist applies least privilege by exposing only the file-analysis capabilities needed for this task. The investigator cannot silently turn analysis into a repository change.

  • In the file sidebar of the editor you used for incident_router.py, create a .claude folder at the top level of the activated workspace.
  • Inside .claude, create an agents folder.
  • Confirm the file sidebar now shows .claude/agents/.
  • Inside .claude/agents/, create application-investigator.md.

You should see the empty application-investigator.md file open in your editor.

  • Define the application investigator by pasting this code into the open file:
---
name: application-investigator
description: Investigates application behavior and request-handling evidence after the incident router selects the application route.
tools: Read, Grep, Glob
permissionMode: plan
---

You are a read-only application incident investigator. Inspect only the files named in the task. Do not modify files, run commands, or recommend executing a change during the investigation.

Return:
1. Evidence with exact file paths.
2. The most likely application-level cause.
3. Contradicting evidence.
4. Safe next checks for a human operator.

How does this stay read-only?

  • The YAML frontmatter names the agent so the router can select it consistently.
  • The tools: Read, Grep, Glob allowlist exposes file inspection without file modification or command execution.
  • The permissionMode: plan setting keeps the session focused on read-only exploration.
  • The response contract requires exact evidence paths so a human operator can verify every conclusion.
  • Save .claude/agents/application-investigator.md.
  • Confirm the saved file contains the application role plus four numbered report requirements.

Application agent missing?

Check that the file sits directly inside .claude/agents/. A different folder location prevents Claude Code from treating it as a project agent.

Confirm the filename is exactly application-investigator.md. The name in the frontmatter must also remain application-investigator.

Help me check my application investigator definition.

Add deployment and platform scopes

The remaining agents share the same access boundary. Their prompts change the kind of evidence they prioritize.

  • Inside .claude/agents/, create deployment-investigator.md.

You should see deployment-investigator.md beside the application investigator in the file sidebar.

  • Define the deployment investigator by pasting this code into the open file:
---
name: deployment-investigator
description: Investigates release and configuration evidence after the incident router selects the deployment route.
tools: Read, Grep, Glob
permissionMode: plan
---

You are a read-only deployment incident investigator. Inspect only the files named in the task. Do not modify files, run commands, or recommend executing a change during the investigation.

Return:
1. Evidence with exact file paths.
2. The most likely release or configuration cause.
3. Contradicting evidence.
4. Safe next checks for a human operator.

What changes for deployment?

This role examines release evidence or configuration evidence. Its tool boundary remains identical to the application agent.

The response contract asks for a likely release or configuration cause. Contradicting evidence keeps the report from presenting one explanation as certain.

  • Save .claude/agents/deployment-investigator.md.
  • Confirm its frontmatter shows tools: Read, Grep, Glob plus permissionMode: plan.

Deployment definition looks different?

Compare the frontmatter against the application agent. Only the agent name plus its application-specific wording should change.

Check the opening and closing --- lines. They separate the frontmatter from the investigation prompt.

Help me compare my deployment investigator definition.

  • Inside .claude/agents/, create platform-investigator.md.

You should now see three investigator files inside .claude/agents/.

  • Define the platform investigator by pasting this code into the open file:
---
name: platform-investigator
description: Investigates infrastructure, networking, capacity, and shared-platform evidence after the incident router selects the platform route.
tools: Read, Grep, Glob
permissionMode: plan
---

You are a read-only platform incident investigator. Inspect only the files named in the task. Do not modify files, run commands, or recommend executing a change during the investigation.

Return:
1. Evidence with exact file paths.
2. The most likely platform-level cause.
3. Contradicting evidence.
4. Safe next checks for a human operator.

What changes for platform?

This role looks for infrastructure, networking, capacity, or shared-platform evidence. The specialized scope gives the router a focused destination for an approved platform route.

The same four-part report keeps its findings auditable. A human operator receives evidence before any operational change is considered.

  • Save .claude/agents/platform-investigator.md.
  • Confirm all three files appear directly inside .claude/agents/.

You have the access boundary in place. Every investigator now uses the same read-only tool surface.

Platform agent not listed?

Check that platform-investigator.md is saved inside the same .claude/agents/ folder as the other two definitions.

Confirm the frontmatter name is exactly platform-investigator. Claude Code uses that name when selecting the agent.

Help me troubleshoot my platform investigator definition.

✔️ Awesome, I've got everything!

Your three investigator definitions are saved inside .claude/agents/.

ⓧ I'd like to double check the full code

Use these complete files to compare every line against your three investigator definitions.

---
name: application-investigator
description: Investigates application behavior and request-handling evidence after the incident router selects the application route.
tools: Read, Grep, Glob
permissionMode: plan
---

You are a read-only application incident investigator. Inspect only the files named in the task. Do not modify files, run commands, or recommend executing a change during the investigation.

Return:
1. Evidence with exact file paths.
2. The most likely application-level cause.
3. Contradicting evidence.
4. Safe next checks for a human operator.

Application agent reference

This file limits application investigations to named evidence files. It returns a path-cited report for human review.

---
name: deployment-investigator
description: Investigates release and configuration evidence after the incident router selects the deployment route.
tools: Read, Grep, Glob
permissionMode: plan
---

You are a read-only deployment incident investigator. Inspect only the files named in the task. Do not modify files, run commands, or recommend executing a change during the investigation.

Return:
1. Evidence with exact file paths.
2. The most likely release or configuration cause.
3. Contradicting evidence.
4. Safe next checks for a human operator.

Deployment agent reference

This file focuses the same read-only workflow on release or configuration evidence. Its report preserves uncertainty through contradicting evidence.

---
name: platform-investigator
description: Investigates infrastructure, networking, capacity, and shared-platform evidence after the incident router selects the platform route.
tools: Read, Grep, Glob
permissionMode: plan
---

You are a read-only platform incident investigator. Inspect only the files named in the task. Do not modify files, run commands, or recommend executing a change during the investigation.

Return:
1. Evidence with exact file paths.
2. The most likely platform-level cause.
3. Contradicting evidence.
4. Safe next checks for a human operator.

Platform agent reference

This file scopes platform analysis to infrastructure and shared-service evidence. Its plan-mode tool boundary prevents operational changes.

Run the approved platform investigator

The router maps an approved platform route to platform-investigator. You can now test that handoff against the existing cache-latency incident.

Before you run this, which facts from the incident do you expect the investigator to cite?

  • Return to the activated terminal from earlier.
  • Start the approved platform investigation by running this command:
claude --agent platform-investigator --permission-mode plan "Investigate incidents/cache_latency.json. Return evidence, likely cause, and next safe checks. Do not modify files."

What does this launch?

  • The --agent platform-investigator option selects the project-scoped platform role.
  • The --permission-mode plan option begins the session in read-only exploration mode.
  • The prompt limits the investigation to incidents/cache_latency.json.
  • The requested report separates evidence from the likely cause plus the next checks for a human operator.

You should receive a report that cites incidents/cache_latency.json. The report should include evidence, a likely platform-level cause, contradicting evidence, and safe next checks.

Agent does not start?

Confirm the terminal is still inside the activated workspace. Project-scoped agents are discovered from the workspace's .claude/agents/ folder.

Check that the selected name matches platform-investigator in both the command and the file frontmatter.

Help me start the read-only platform investigator.

  • Compare the report's cited statements against incidents/cache_latency.json.
  • Confirm every cited claim appears in the named incident file.

That is the safe handoff working. An approved route now reaches a specialist that can inspect evidence without changing the repository.

Your router now hands approved incidents to constrained specialists. Next, you will turn recorded Jev responses into repeatable regression evidence.

Build a Reproducible Evaluation Suite

Your router now fails closed when Jev is unavailable or unsafe. Approved routes can reach investigators that only inspect evidence.

Live Jev decisions can vary between requests. Network access can also disappear during an incident.

Recorded responses use dependency injection to make every policy branch repeatable. A pytest suite turns those branches into evidence without contacting OpenRouter.

In this step, get ready to:
  • Load every recorded response through one reusable test fixture.
  • Add fixed normal-route and high-risk Jev responses.
  • Prove that safe outcomes dispatch while uncertain or broken outcomes fail closed.
Build the recorded-response loader

A pytest fixture can load every response once for the entire test session. Each test receives the same local payload through recorded_responses.

  • Create tests/conftest.py in the project workspace from earlier by copying this code into your editor:
import json
from pathlib import Path

import pytest


@pytest.fixture(scope="session")
def recorded_responses() -> dict[str, dict]:
    fixture_dir = Path(__file__).parent / "fixtures"
    return {
        path.stem: json.loads(path.read_text(encoding="utf-8"))
        for path in sorted(fixture_dir.glob("*.json"))
    }

What Does This Loader Do?

  • The session scope loads the fixture collection once for the entire test run.
  • The glob("*.json") expression finds every recorded response in tests/fixtures.
  • Each file stem becomes a dictionary key such as normal or low_confidence.
  • The loader keeps the complete payload available for policy checks or audit inspection.
  • Save tests/conftest.py.
  • Create tests/test_policy.py with the unavailable-provider test by copying this code into your editor:
import pytest

from incident_router import route_incident


def test_unavailable_jev_escalates():
    def unavailable(_incident):
        raise ConnectionError("simulated outage")

    result = route_incident({}, provider=unavailable)

    assert result.action == "human_review"
    assert result.reason == "jev_unavailable"
    assert result.agent is None

Why Inject a Failing Provider?

The local unavailable function simulates a provider outage. Passing it into route_incident() exercises the same boundary used by the live provider.

The assertions prove that an exception produces human review with no agent dispatch. The test never attempts a network request.

  • Save tests/test_policy.py.
  • Prove the unavailable-provider guardrail works by running:
python3 -m pytest -q

What Does This Test Command Do?

Python runs the installed pytest module against the project. The quiet option keeps the report focused on test outcomes.

You should see one passing test. That first guardrail is now proven: a provider outage reaches human review without dispatching an investigator.

Is the First Test Failing?

Confirm that tests/conftest.py and tests/test_policy.py are inside the existing project workspace.

Check that your virtual environment from Step 1 remains active. The pinned pytest package must be available to this Python process.

Help me debug the unavailable-provider pytest test.

Add fixed Jev response fixtures

Recorded fixtures freeze the exact payload presented to Python. The normal case approves a platform investigator with strong confidence.

  • Create tests/fixtures/normal.json by copying this response into your editor:
{
  "fixture_source": "authored_test_double",
  "model": "typesafe/jev-1.13-20260917",
  "answers": {
    "investigator": {
      "type": "choice",
      "choice": "platform",
      "confidence": 0.93,
      "probabilities": {
        "application": 0.04,
        "deployment": 0.03,
        "platform": 0.93
      }
    },
    "human_review": {
      "type": "noul",
      "noul": 0.08
    }
  },
  "usage": {
    "input_tokens": 310,
    "output_tokens": 24,
    "cost": 0.000013
  }
}

What Does the Normal Fixture Capture?

  • The serving model belongs to the pinned Jev model family.
  • The selected route is platform with confidence 0.93.
  • The human-review probability is 0.08.
  • The usage object preserves token counts and reported cost for later inspection.
  • Save tests/fixtures/normal.json.
  • Validate the normal fixture by running:
python3 -m json.tool tests/fixtures/normal.json

What Does This Validation Show?

Python parses the fixture as JSON before printing a formatted copy. Reaching the formatted output confirms that the file has valid JSON syntax.

You should see the complete normal payload formatted in your terminal. Its audit fields remain available without replaying a live request.

The high-risk fixture tests a different safety boundary. Its route confidence looks strong while its human-review probability requires escalation.

  • Create tests/fixtures/high_risk.json by copying this response into your editor:
{
  "fixture_source": "authored_test_double",
  "model": "typesafe/jev-1.13-20260917",
  "answers": {
    "investigator": {
      "type": "choice",
      "choice": "application",
      "confidence": 0.91,
      "probabilities": {
        "application": 0.91,
        "deployment": 0.05,
        "platform": 0.04
      }
    },
    "human_review": {
      "type": "noul",
      "noul": 0.82
    }
  },
  "usage": {
    "input_tokens": 320,
    "output_tokens": 24,
    "cost": 0.000014
  }
}

What Makes This Fixture High Risk?

Jev selects the application route with confidence 0.91. The separate human-review probability is 0.82.

Python evaluates both signals before dispatch. The high review probability must override the confident route.

  • Save tests/fixtures/high_risk.json.
  • Validate the high-risk fixture by running:
python3 -m json.tool tests/fixtures/high_risk.json

What Does This Validation Prove?

The command parses the second fixture independently. A formatted payload confirms that pytest can load it as JSON.

You should see the complete high-risk payload in the terminal. Both new fixtures are now valid recorded inputs.

Is a Fixture Failing to Parse?

Compare the affected file with its code block. Check every comma after a property or nested object.

Confirm that every property name uses double quotes. JSON does not accept Python-style single-quoted strings.

Help me find the JSON syntax problem in my recorded fixture.

Parametrize every policy outcome

Parametrization runs one test body against several recorded outcomes. Injecting each payload into route_incident() isolates deterministic policy from the live model service.

  • In tests/test_policy.py place this test between the imports and test_unavailable_jev_escalates():
@pytest.mark.parametrize(
    "fixture_name,expected_action,expected_reason,expected_agent",
    [
        (
            "normal",
            "dispatch_read_only_investigator",
            "policy_approved",
            "platform-investigator",
        ),
        ("low_confidence", "human_review", "low_route_confidence", None),
        (
            "high_risk",
            "human_review",
            "high_human_review_probability",
            None,
        ),
        ("invalid", "human_review", "invalid_jev_response", None),
    ],
)
def test_recorded_policy_outcomes(
    recorded_responses,
    fixture_name,
    expected_action,
    expected_reason,
    expected_agent,
):
    payload = recorded_responses[fixture_name]
    result = route_incident({}, provider=lambda _incident: payload)

    assert result.action == expected_action

How Does the Parametrized Test Work?

  • The parameter table defines normal, low-confidence, high-risk, and invalid-response cases.
  • The recorded_responses fixture supplies the selected local payload.
  • The injected lambda returns that payload without contacting OpenRouter.
  • The first assertion checks whether the router dispatches or escalates.
  • Save tests/test_policy.py.
  • Run the expanded action checks with:
python3 -m pytest -q

What Does This Run Cover?

Pytest runs four recorded cases plus the unavailable-provider case. Every provider is local or deliberately raises an exception.

You should see five passing tests. The suite now distinguishes dispatch from escalation without making a network request.

Action alone cannot distinguish one escalation rule from another. Two more assertions pin the exact reason and agent outcome.

  • In tests/test_policy.py find this assertion:
    assert result.action == expected_action

What Does the Current Assertion Check?

This assertion checks the top-level action for each recorded response. It does not yet verify which rule produced that action.

  • Replace that assertion with this complete assertion group:
    assert result.action == expected_action
    assert result.reason == expected_reason
    assert result.agent == expected_agent

What Do the Final Assertions Prove?

The reason assertion ties each unsafe payload to its expected fail-closed rule. The agent assertion proves that only the approved normal route selects platform-investigator.

  • Save tests/test_policy.py.

Before you run the final check, which recorded case do you expect to be the only one that dispatches an investigator?

  • Run the complete deterministic evaluation suite with:
python3 -m pytest -q

What Does the Final Run Prove?

The normal fixture dispatches platform-investigator at route confidence 0.93. Its recorded model, token usage, and cost remain available in the loaded payload.

Low confidence, high review probability, invalid data, and provider failure all reach human review. The same fixture produces the same policy result on every run.

You should see five passing tests. You now have repeatable regression evidence that every unsafe or broken model outcome fails closed.

Is the Full Suite Failing?

Use the failing case name to identify the matching file in tests/fixtures. Compare its values with the expected row in tests/test_policy.py.

If fixture loading fails, confirm that every JSON file in tests/fixtures contains one object at its root.

Help me debug my deterministic incident-router test suite.

✔️ Awesome, I've got everything!

Great. Confirm that you saved every new fixture and both Python test files.

ⓧ I'd like to double check the full code

Compare your complete tests/conftest.py file with this version.

import json
from pathlib import Path

import pytest


@pytest.fixture(scope="session")
def recorded_responses() -> dict[str, dict]:
    fixture_dir = Path(__file__).parent / "fixtures"
    return {
        path.stem: json.loads(path.read_text(encoding="utf-8"))
        for path in sorted(fixture_dir.glob("*.json"))
    }

This file loads every JSON fixture into the session-scoped recorded_responses dictionary.

Compare your complete tests/fixtures/normal.json file with this version.

{
  "fixture_source": "authored_test_double",
  "model": "typesafe/jev-1.13-20260917",
  "answers": {
    "investigator": {
      "type": "choice",
      "choice": "platform",
      "confidence": 0.93,
      "probabilities": {
        "application": 0.04,
        "deployment": 0.03,
        "platform": 0.93
      }
    },
    "human_review": {
      "type": "noul",
      "noul": 0.08
    }
  },
  "usage": {
    "input_tokens": 310,
    "output_tokens": 24,
    "cost": 0.000013
  }
}

This fixture records the policy-approved platform route with its complete audit metadata.

Compare your complete tests/fixtures/high_risk.json file with this version.

{
  "fixture_source": "authored_test_double",
  "model": "typesafe/jev-1.13-20260917",
  "answers": {
    "investigator": {
      "type": "choice",
      "choice": "application",
      "confidence": 0.91,
      "probabilities": {
        "application": 0.91,
        "deployment": 0.05,
        "platform": 0.04
      }
    },
    "human_review": {
      "type": "noul",
      "noul": 0.82
    }
  },
  "usage": {
    "input_tokens": 320,
    "output_tokens": 24,
    "cost": 0.000014
  }
}

This fixture records a confident application route whose human-review probability still forces escalation.

Compare your complete tests/test_policy.py file with this version.

import pytest

from incident_router import route_incident


@pytest.mark.parametrize(
    "fixture_name,expected_action,expected_reason,expected_agent",
    [
        (
            "normal",
            "dispatch_read_only_investigator",
            "policy_approved",
            "platform-investigator",
        ),
        ("low_confidence", "human_review", "low_route_confidence", None),
        (
            "high_risk",
            "human_review",
            "high_human_review_probability",
            None,
        ),
        ("invalid", "human_review", "invalid_jev_response", None),
    ],
)
def test_recorded_policy_outcomes(
    recorded_responses,
    fixture_name,
    expected_action,
    expected_reason,
    expected_agent,
):
    payload = recorded_responses[fixture_name]
    result = route_incident({}, provider=lambda _incident: payload)

    assert result.action == expected_action
    assert result.reason == expected_reason
    assert result.agent == expected_agent


def test_unavailable_jev_escalates():
    def unavailable(_incident):
        raise ConnectionError("simulated outage")

    result = route_incident({}, provider=unavailable)

    assert result.action == "human_review"
    assert result.reason == "jev_unavailable"
    assert result.agent is None

This file injects recorded responses for four policy outcomes. It also simulates an unavailable provider without contacting OpenRouter.

Secret mission

Hidden Authentication Outage

Build an adversarial incident where a confident Jev deployment route collides with an explicit authentication-outage flag. Add a deterministic regression test that proves Python escalates to human review without dispatching an investigator.

Clean Up Your Resources

Clean Up Your Resources

The local files create no ongoing charges. Decide whether to keep them ready, pause your work, or delete the router artifacts.

Resources you used:

  • The .venv Python environment plus the generated __pycache__/ and .pytest_cache/ cache folders.
  • The router source and configuration files: incident_router.py, requirements.txt, and .gitignore.
  • The incident inputs in incidents/cache_latency.json and incidents/hidden_auth_outage.json.
  • The deterministic evaluation suite in tests/, including the recorded response fixtures and policy tests.
  • The three Claude Code agent definitions in .claude/agents/: application-investigator.md, deployment-investigator.md, and platform-investigator.md.

Keep everything running

No action is needed. Choose this if you want to keep testing the router or extend its safety policy.

  • Keep all router artifacts in the current workspace.
  • Replay the recorded fixtures when you want deterministic results without making a network request.
  • Run the router live only when you want a fresh Jev decision through OpenRouter that may consume account credit.

Pause - I'll come back to this later

There are no servers or background router processes to stop. Choose this if you want to leave the workspace on disk for another session.

  • Close the Terminal session you used for the project.
  • Keep .venv in the workspace so the pinned dependencies remain available.
  • Keep tests/fixtures/live_cache_latency.json if you recorded a live response that you want to audit or replay later.

Delete - I don't want to use this again

Deleting the router artifacts is permanent. The removal command names each target explicitly so unrelated files stay in place.

  • Switch back to the Terminal session from earlier.
  • Confirm that the Terminal is in the project workspace by running this command:
ls

What should I see?

You should see incident_router.py, requirements.txt, incidents, and tests in the output. These names confirm that the Terminal is in the project workspace.

  • Delete the local environment plus every router artifact by running this command:
rm -rf .venv __pycache__ .pytest_cache .claude incidents tests .gitignore incident_router.py requirements.txt

What does this remove?

This command removes the local Python environment plus the named router files. It also removes the incidents, fixtures, tests, caches, and project-scoped agents.

Your existing OpenRouter account remains in place. Your Claude Code installation also remains available.

  • Verify that the visible router artifacts are gone by running this command:
ls

What should I see now?

You should no longer see incident_router.py, requirements.txt, incidents, or tests in the output. That confirms the router artifacts have been removed.

Nice Work!

Nice Work!

Excellent work. You've built a fail-safe incident router that places deterministic Python policy above every Jev recommendation.

You've learned how to:

  • Sent a live incident decision through OpenRouter using the TypeSafe Python SDK with jev-1.13. Inspected the selected route, confidence, serving model, token usage, and cost when OpenRouter reports it.
  • Made deterministic Python policy the final authority over model output. Unavailable providers, invalid responses, elevated human-review probability, and route confidence below 0.80 now fail closed to human_review.
  • Restricted three project-scoped Claude Code investigators to Read, Grep, and Glob in plan mode. Backed those safe dispatches with pytest regression tests that inject recorded responses without contacting OpenRouter.
  • Secret Mission: Proved an authentication-outage policy overrides a confident deployment recommendation. The hidden case still returns human_review at 0.97 route confidence with no agent dispatch.

Ready to quiz yourself?