Build a Resumable LangGraph Workflow

Build a resumable incident workflow with LangGraph, approvals, and audit trails.

Introduction

30 Second Summary

When a live service fails, every decision happens under pressure. An automatic fix can deepen the damage when nobody reviews it.

In this project, you will build a command-line incident-triage workflow in Google Cloud Shell. Gemini 3.7 Flash produces structured incident analysis while LangGraph holds high-risk plans for human approval.

What You'll Build

You will demo a risky incident pausing for approval before a new Python process resumes its saved state to publish an approved runbook.

By the end of this project, you'll have:

  • An unsafe-versus-protected comparison where the naive workflow publishes immediately while the gated workflow stops with an approval payload.
  • A thread-scoped approval flow that loads the exact incident from a SQLite checkpoint when you approve or reject it in a separate command.
  • An auditable Markdown runbook that shows the original report, structured analysis, policy result, human decision, and ordered audit log.
  • Secret Mission: Add a checkpoint inspection command that reveals saved analysis, workflow status, checkpoint time, and pending nodes without resuming the workflow.

Are there any prerequisites?

You need an existing Google Cloud project with billing enabled. Your account also needs permission to enable the Gemini Enterprise Agent Platform API.

Before We Start

An AI-generated high-risk production remediation plan can cause harm if a workflow publishes it before an operator reviews it. This checkpoint commits you to preserving that human accountability boundary before any hands-on work begins.

Set Up the Cloud Shell Environment

Your incident workflow needs a reproducible environment before it can analyze reports or save approval checkpoints. Google Cloud Shell removes the local Mac setup while keeping your project files in persistent storage.

A virtual environment isolates the workflow packages from other Python projects. The enabled model API lets later steps access Gemini through Cloud Shell's existing Application Default Credentials.

In this step, get ready to:
  • Confirm the billed Google Cloud project used by Cloud Shell.
  • Enable the Agent Platform API for that project.
  • Prepare an isolated project directory with the pinned Python packages.
Confirm project and API access

Cloud Shell sets GOOGLE_CLOUD_PROJECT to the active Google Cloud project. Confirming this value prevents you from enabling services in the wrong billed project.

  • Sign in to the Google Cloud console with the account that owns your billed project.
  • Select your billed project from the project selector in the console header.
  • Click Activate Cloud Shell at the top of the console.

Cloud Shell can take a few seconds to initialize. You are ready when the terminal displays a command prompt.

  • Check the project attached to your Cloud Shell session by running this command:
echo "$GOOGLE_CLOUD_PROJECT"

What does this command show?

The command prints the project ID stored in GOOGLE_CLOUD_PROJECT. Google client libraries use this active project with Cloud Shell's existing credentials.

  • Confirm the printed project ID matches the billed project you selected.

The Gemini Enterprise Agent Platform API allows the workflow to make model requests in later steps. Enabling aiplatform.googleapis.com does not create an always-running service.

  • Enable the API in the active project by running this command:
gcloud services enable aiplatform.googleapis.com

What does this command do?

The gcloud command enables the model API for the project identified by your active Cloud Shell configuration. Later model requests depend on this service being available.

The command returns to the prompt without an error when the API is enabled.

Unable to enable the API?

Confirm that the printed project ID belongs to the project where you have permission to enable services. An organization policy can also block service activation.

If access is denied, ask the project owner for permission to enable services before continuing.

Help me troubleshoot Agent Platform API access.

Prepare the isolated project

The pinned packages require a current Python 3 runtime. Cloud Shell provides Python 3.12, which meets their supported version requirements.

  • Check the Python version provided by Cloud Shell by running this command:
python3 --version

What does this command show?

The version check confirms which Python interpreter creates your virtual environment. The expected Cloud Shell runtime is Python 3.12.

✔️ I see version 3.12

Your Cloud Shell Python runtime is ready for the pinned packages. Continue with the project directory setup below.

ⓧ I see an older version

Cloud Shell officially provides Python 3.12. An older result suggests that the current virtual machine needs to be restarted.

  • Click More in the Cloud Shell toolbar.
  • Click Restart.
  • Wait for the new Cloud Shell prompt to load.
  • Repeat the Python version check above.

Still seeing an older version?

Stop before installing the dependencies because an older interpreter can reject the pinned packages.

Help me restore the current Cloud Shell Python runtime.

ⓧ Command not found

Python 3 is preinstalled in Cloud Shell. A missing command suggests that the current virtual machine or shell configuration needs a fresh session.

  • Click More in the Cloud Shell toolbar.
  • Click Restart.
  • Wait for the new Cloud Shell prompt to load.
  • Repeat the Python version check above.

Python still unavailable?

A customized shell path can hide preinstalled commands even after the virtual machine restarts.

Help me restore Python in Cloud Shell.

The project directory keeps your workflow code in Cloud Shell's persistent home storage. The .venv directory keeps this project's packages separate from the managed system installation.

  • Create the project directory with its virtual environment by running these commands:
mkdir -p ~/langgraph-incident-triage
cd ~/langgraph-incident-triage
python3 -m venv .venv
source .venv/bin/activate

What do these commands create?

  • The first command creates ~/langgraph-incident-triage inside your persistent home storage.
  • The second command makes that directory your active terminal location.
  • The third command creates an isolated Python environment in .venv.
  • The final command activates that environment for the current terminal session.
  • Check that your terminal prompt starts with (.venv).
  • Check that the prompt shows ~/langgraph-incident-triage as the current directory.

Virtual environment not active?

Confirm that you ran the activation command from inside ~/langgraph-incident-triage. A missing (.venv) prefix means the project is still using the system Python environment.

Help me activate this virtual environment.

The dependency file pins LangGraph, the SQLite checkpointer, and LangChain Google GenAI to the versions used throughout this project.

  • Click Open Editor in the Cloud Shell toolbar.
  • Select the langgraph-incident-triage folder in the editor file sidebar.
  • Click the new file icon beside the selected folder.
  • Name the file requirements.txt.
  • Paste the pinned dependencies below into the new file:
langgraph==1.2.14
langgraph-checkpoint-sqlite==3.1.1
langchain-google-genai==4.4.0

Why pin these versions?

Exact version pins make the environment reproducible. Every learner uses the same graph API, checkpoint implementation, and model integration.

  • Save requirements.txt.
  • Confirm the editor shows exactly three dependency lines.

File saved in the wrong place?

Check that requirements.txt appears directly inside ~/langgraph-incident-triage. It should sit beside the .venv directory.

Help me place the dependency file correctly.

✔️ Awesome, I've got everything!

Great. Keep requirements.txt saved before returning to the terminal.

ⓧ I'd like to double check the full code

langgraph==1.2.14
langgraph-checkpoint-sqlite==3.1.1
langchain-google-genai==4.4.0
Finish the Python environment

The activated virtual environment now has its own package installer. Installing from requirements.txt reproduces the exact package set used by the workflow code.

  • Switch back to the Cloud Shell terminal.
  • Upgrade the environment's package installer with the first command below.
  • Install the pinned project dependencies with the second command below:
python3 -m pip install --upgrade pip
python3 -m pip install -r requirements.txt

What do these commands install?

The first command updates pip inside the active virtual environment. The second command reads each exact package version from requirements.txt.

Package installation can take a few minutes while Cloud Shell downloads the dependencies. A stream of package messages means the installation is still running.

  • Wait until the terminal returns to the (.venv) prompt.

Package installation failed?

Confirm that the prompt starts with (.venv). Also compare every line in requirements.txt with the double-check tab above.

A temporary network interruption can stop a package download. Repeat the installation after Cloud Shell reconnects.

Help me troubleshoot the dependency installation.

Before you run the final check, which three imports do you expect the environment to load successfully?

  • Verify all three project libraries with this import check:
python3 -c "from langgraph.graph import StateGraph; from langgraph.checkpoint.sqlite import SqliteSaver; from langchain_google_genai import ChatGoogleGenerativeAI; print('LangGraph environment ready')"

What does this check prove?

The command imports the graph builder, the SQLite checkpointer, and the Gemini chat integration. The final print runs only after every import succeeds.

You should see LangGraph environment ready in the terminal. This confirms that the isolated environment can load every library needed by the workflow.

Import check not passing?

Confirm that your terminal prompt still starts with (.venv). A new Cloud Shell terminal does not automatically reactivate the environment.

Compare the failed import name with the three package lines in requirements.txt.

Help me fix the failing import.

That's the foundation in place: your Cloud Shell project can now load the graph, checkpoint, and Gemini integration libraries. Next, you will build a working graph that exposes the risk of publishing an AI-generated runbook without review.

Build the Unsafe Auto-Publisher

Your activated Cloud Shell environment is ready to run a LangGraph workflow. You can now turn a raw incident report into a concrete runbook.

A direct workflow lets Gemini 3.7 Flash analysis reach the publishing step without an approval boundary. This step deliberately builds that unsafe path so you can see the control gap for yourself.

In this step, get ready to:
  • Define structured incident analysis for the sample production incident.
  • Connect analysis directly to runbook publishing.
  • Run the workflow to expose the missing approval control.
Define structured incident analysis

Structured model output gives each incident analysis a predictable shape. The workflow can then read fields such as severity or risk without parsing free-form prose.

  • Open the Cloud Shell editor from the Cloud Shell toolbar.
  • Create naive_app.py inside ~/langgraph-incident-triage using the editor's file creation control.
  • Add the imports plus the sample incident by pasting this code:
import json
import os
from pathlib import Path
from typing import TypedDict

from langchain_google_genai import ChatGoogleGenerativeAI
from langgraph.graph import END, START, StateGraph

MODEL_ID = "gemini-3.7-flash"
INCIDENT = (
    "Production checkout is failing for all users. "
    "The proposed emergency action is to restart the production database."
)

What Does This Code Set Up?

  • The imports provide JSON formatting plus file access plus typed state plus the model integration plus graph construction.
  • The MODEL_ID constant selects gemini-3.7-flash for incident analysis.
  • The INCIDENT constant contains synthetic production symptoms plus a proposed operational action.
  • Save naive_app.py.
  • Confirm the editor shows MODEL_ID above INCIDENT.

Imports or Constants Look Incomplete?

Check that naive_app.py sits directly inside ~/langgraph-incident-triage. Make sure both incident strings remain inside the parentheses.

Ask for help reviewing this first code section:

A JSON Schema defines the allowed fields plus values for the model response. This contract keeps the rest of the workflow independent from variations in model wording.

  • Place your cursor below the closing parenthesis for INCIDENT.
  • Add the analysis schema by pasting this code:
ANALYSIS_SCHEMA = {
    "title": "IncidentAnalysis",
    "type": "object",
    "properties": {
        "severity": {
            "type": "string",
            "enum": ["SEV1", "SEV2", "SEV3", "SEV4"],
        },
        "summary": {"type": "string"},
        "proposed_action": {"type": "string"},
        "risk": {
            "type": "string",
            "enum": ["low", "high"],
        },
        "rationale": {"type": "string"},
    },
    "required": [
        "severity",
        "summary",
        "proposed_action",
        "risk",
        "rationale",
    ],
}

How Does the Schema Control the Response?

  • The severity field must use one of four incident levels.
  • The risk field must use either low or high.
  • The required list makes every analysis field mandatory.
  • The stable structure allows later graph nodes to consume the result as data.
  • Save naive_app.py.
  • Confirm the schema ends with the five required field names.

Seeing Unmatched Brackets?

Compare each opening brace with its closing brace. Check that every field except the final field in each section has a trailing comma.

Ask for help finding the syntax mismatch:

The graph state carries data between nodes. The analyst factory connects to Vertex AI through the active Cloud Shell project.

  • Place your cursor below ANALYSIS_SCHEMA.
  • Add the state definition plus analyst factory by pasting this code:
class NaiveState(TypedDict, total=False):
    incident: str
    analysis: dict[str, str]
    status: str
    runbook_path: str


def build_analyst():
    project_id = os.environ.get("GOOGLE_CLOUD_PROJECT")
    if not project_id:
        raise RuntimeError("GOOGLE_CLOUD_PROJECT is not set.")

    model = ChatGoogleGenerativeAI(
        model=MODEL_ID,
        project=project_id,
        location="global",
        vertexai=True,
        max_retries=2,
    )
    return model.with_structured_output(
        schema=ANALYSIS_SCHEMA,
        method="json_schema",
    )

How Does the Analyst Use Your Environment?

  • The NaiveState definition lists the values that graph nodes can exchange.
  • The build_analyst() function reads the active project from GOOGLE_CLOUD_PROJECT.
  • The model configuration uses the global location plus the Vertex AI backend.
  • The structured-output wrapper applies ANALYSIS_SCHEMA to every incident response.
  • Save naive_app.py.
  • Confirm build_analyst() ends with the structured-output configuration.

Project Configuration Look Wrong?

Make sure the environment variable name is exactly GOOGLE_CLOUD_PROJECT. Check that project_id is passed to the model's project parameter.

Ask for help checking the model setup:

The analysis node is the workflow's model-powered step. It sends the incident to the analyst before returning structured data to the graph.

  • Place your cursor below build_analyst().
  • Add the incident analysis node by pasting this code:
def analyze_incident(state: NaiveState):
    analyst = build_analyst()
    analysis = analyst.invoke(
        [
            (
                "system",
                "You are an SRE incident analyst. Classify the incident and "
                "propose one reversible remediation step. Keep every field concise.",
            ),
            ("human", state["incident"]),
        ]
    )
    return {"analysis": analysis, "status": "analyzed"}

What Does the Analysis Node Do?

  • The system message gives the model an incident-analysis role.
  • The human message supplies the current state's incident value.
  • The returned update stores the structured model result under analysis.
  • The analyzed status records that the model step completed.
  • Save naive_app.py.
  • Confirm analyze_incident() returns both analysis plus status.

Analysis Node Indentation Look Uneven?

Keep the two message tuples inside the list passed to analyst.invoke(). Align the closing list bracket with the opening call.

Ask for help checking the node structure:

Connect the direct publishing path

A graph node becomes consequential when it performs a side effect. This publisher writes the model proposal to a Markdown file on disk.

  • Place your cursor below analyze_incident().
  • Add the runbook publisher by pasting this code:
def publish_runbook(state: NaiveState):
    analysis = state["analysis"]
    path = Path("naive_runbook.md")
    path.write_text(
        "# Unreviewed Incident Runbook\n\n"
        f"Incident: {state['incident']}\n\n"
        f"Severity: {analysis['severity']}\n\n"
        f"Summary: {analysis['summary']}\n\n"
        f"Proposed action: {analysis['proposed_action']}\n\n"
        f"Risk: {analysis['risk']}\n\n"
        "Approval: Not requested\n",
        encoding="utf-8",
    )
    return {
        "status": "published_without_review",
        "runbook_path": str(path),
    }

What Makes Publishing a Side Effect?

  • The function reads the structured model analysis from the graph state.
  • The write_text() call creates naive_runbook.md on disk.
  • The runbook records that approval was never requested.
  • The returned published_without_review status makes the unsafe outcome visible in the terminal.
  • Save naive_app.py.
  • Confirm the publisher targets naive_runbook.md.

Runbook Strings Look Broken?

Keep each runbook line inside its matching quotes. Check that every newline sequence remains inside the string.

Ask for help reviewing the file-writing code:

The final function assembles the graph. Its edges place publishing immediately after model analysis.

  • Place your cursor below publish_runbook().
  • Add the direct graph plus script entry point by pasting this code:
def main():
    builder = StateGraph(NaiveState)
    builder.add_node("analyze_incident", analyze_incident)
    builder.add_node("publish_runbook", publish_runbook)
    builder.add_edge(START, "analyze_incident")
    builder.add_edge("analyze_incident", "publish_runbook")
    builder.add_edge("publish_runbook", END)

    graph = builder.compile()
    result = graph.invoke({"incident": INCIDENT})
    print(json.dumps(result, indent=2))


if __name__ == "__main__":
    main()

How Does the Direct Graph Flow?

  • The StateGraph uses NaiveState as its shared data contract.
  • The first edge routes START to analyze_incident.
  • The next edge routes analysis directly to publish_runbook.
  • The final edge ends the workflow immediately after publication.
  • Save naive_app.py.
  • Confirm the graph contains a direct edge from analyze_incident to publish_runbook.

Graph Names Do Not Match?

Check that each node name matches its function name exactly. Verify that START plus END use uppercase letters.

Ask for help tracing the graph:

✔️ Awesome, I've got everything!

Great. Double-check that naive_app.py is saved before you run the workflow.

ⓧ I'd like to double check the full code

import json
import os
from pathlib import Path
from typing import TypedDict

from langchain_google_genai import ChatGoogleGenerativeAI
from langgraph.graph import END, START, StateGraph

MODEL_ID = "gemini-3.7-flash"
INCIDENT = (
    "Production checkout is failing for all users. "
    "The proposed emergency action is to restart the production database."
)

ANALYSIS_SCHEMA = {
    "title": "IncidentAnalysis",
    "type": "object",
    "properties": {
        "severity": {
            "type": "string",
            "enum": ["SEV1", "SEV2", "SEV3", "SEV4"],
        },
        "summary": {"type": "string"},
        "proposed_action": {"type": "string"},
        "risk": {
            "type": "string",
            "enum": ["low", "high"],
        },
        "rationale": {"type": "string"},
    },
    "required": [
        "severity",
        "summary",
        "proposed_action",
        "risk",
        "rationale",
    ],
}


class NaiveState(TypedDict, total=False):
    incident: str
    analysis: dict[str, str]
    status: str
    runbook_path: str


def build_analyst():
    project_id = os.environ.get("GOOGLE_CLOUD_PROJECT")
    if not project_id:
        raise RuntimeError("GOOGLE_CLOUD_PROJECT is not set.")

    model = ChatGoogleGenerativeAI(
        model=MODEL_ID,
        project=project_id,
        location="global",
        vertexai=True,
        max_retries=2,
    )
    return model.with_structured_output(
        schema=ANALYSIS_SCHEMA,
        method="json_schema",
    )


def analyze_incident(state: NaiveState):
    analyst = build_analyst()
    analysis = analyst.invoke(
        [
            (
                "system",
                "You are an SRE incident analyst. Classify the incident and "
                "propose one reversible remediation step. Keep every field concise.",
            ),
            ("human", state["incident"]),
        ]
    )
    return {"analysis": analysis, "status": "analyzed"}


def publish_runbook(state: NaiveState):
    analysis = state["analysis"]
    path = Path("naive_runbook.md")
    path.write_text(
        "# Unreviewed Incident Runbook\n\n"
        f"Incident: {state['incident']}\n\n"
        f"Severity: {analysis['severity']}\n\n"
        f"Summary: {analysis['summary']}\n\n"
        f"Proposed action: {analysis['proposed_action']}\n\n"
        f"Risk: {analysis['risk']}\n\n"
        "Approval: Not requested\n",
        encoding="utf-8",
    )
    return {
        "status": "published_without_review",
        "runbook_path": str(path),
    }


def main():
    builder = StateGraph(NaiveState)
    builder.add_node("analyze_incident", analyze_incident)
    builder.add_node("publish_runbook", publish_runbook)
    builder.add_edge(START, "analyze_incident")
    builder.add_edge("analyze_incident", "publish_runbook")
    builder.add_edge("publish_runbook", END)

    graph = builder.compile()
    result = graph.invoke({"incident": INCIDENT})
    print(json.dumps(result, indent=2))


if __name__ == "__main__":
    main()
Run the designed failure

The workflow now has everything required to analyze the sample incident plus write a runbook. Running it reveals whether the direct edge offers any chance for operator review.

  • Return to the Cloud Shell terminal from earlier.
  • Make a mental prediction about whether a runbook appears before any operator decision.
  • Run the workflow plus display its generated runbook by entering these commands:
python3 naive_app.py
cat naive_runbook.md

What Should You See?

The model request can take a moment to return. Once it completes, the printed state contains published_without_review plus the path to naive_runbook.md.

The displayed runbook contains a structured proposal. Its approval line says Approval: Not requested.

  • Confirm naive_runbook.md contains the original incident plus the model's severity plus its proposed action.
  • Confirm the terminal reports published_without_review.

This Failure Is Intentional

The graph worked exactly as written. Its direct edge allowed an AI-generated production proposal to become a visible runbook without human review.

That missing boundary is the risk this project is designed to expose. The next workflow closes the gap before any publishing side effect.

Workflow Did Not Produce the Runbook?

Confirm the activated environment still appears in your Cloud Shell prompt. Check that the terminal remains inside ~/langgraph-incident-triage.

If the model request fails, verify that GOOGLE_CLOUD_PROJECT still identifies the billed project from the previous step.

Ask for help diagnosing the run:

You have made the unsafe behavior visible: the direct graph publishes a model-generated runbook without review. Next up, you will add deterministic risk policy plus a durable human approval gate.

Add Risk Policy and Human Approval

Your naive LangGraph workflow proved that a model-generated plan can reach publication without review. That automatic path is dangerous when an incident affects production systems.

This step separates model analysis from deterministic policy. It also adds a human approval boundary before the publishing side effect.

In this step, get ready to:
  • Build a protected graph with explicit incident state and deterministic risk rules.
  • Add a durable approval interrupt backed by a SQLite checkpoint.
  • Prove that a high-risk incident cannot publish a runbook before approval.
Build the protected graph

The protected workflow needs an explicit state that carries the incident through analysis, policy, approval, and publication. The deterministic policy evaluates the model output before any file can be written.

  • Use the Cloud Shell file editor to create app.py inside ~/langgraph-incident-triage.
  • Add the imports, secure checkpoint setting, and workflow constants by pasting this first chunk:
import argparse
import json
import os
from pathlib import Path
from typing import Literal, TypedDict

os.environ.setdefault("LANGGRAPH_STRICT_MSGPACK", "true")

from langchain_google_genai import ChatGoogleGenerativeAI
from langgraph.checkpoint.sqlite import SqliteSaver
from langgraph.graph import END, START, StateGraph
from langgraph.types import Command, interrupt

MODEL_ID = "gemini-3.7-flash"
DB_PATH = "incident_checkpoints.db"
RUNBOOK_DIR = Path("runbooks")
RESTRICTED_TERMS = (
    "delete",
    "disable",
    "drop",
    "production",
    "restart",
    "rotate",
)

What does this setup do?

  • The LANGGRAPH_STRICT_MSGPACK setting restricts checkpoint deserialization to known-safe types.
  • The DB_PATH value gives every process the same checkpoint database.
  • The RESTRICTED_TERMS tuple gives policy enforcement a stable list of operational risk signals.
  • Add the structured analysis schema below the constants by pasting this chunk:
ANALYSIS_SCHEMA = {
    "title": "IncidentAnalysis",
    "type": "object",
    "properties": {
        "severity": {
            "type": "string",
            "enum": ["SEV1", "SEV2", "SEV3", "SEV4"],
        },
        "summary": {"type": "string"},
        "proposed_action": {"type": "string"},
        "risk": {
            "type": "string",
            "enum": ["low", "high"],
        },
        "rationale": {"type": "string"},
    },
    "required": [
        "severity",
        "summary",
        "proposed_action",
        "risk",
        "rationale",
    ],
}

Why keep structured analysis?

The schema constrains the model response to named fields that the graph can evaluate. The policy can check severity and proposed_action without parsing free-form prose.

  • Add the protected state definition and model builder below the schema by pasting this chunk:
class IncidentState(TypedDict, total=False):
    incident_id: str
    incident: str
    analysis: dict[str, str]
    requires_approval: bool
    policy_reason: str
    decision: bool
    status: str
    audit_log: list[str]
    runbook_path: str


def build_analyst():
    project_id = os.environ.get("GOOGLE_CLOUD_PROJECT")
    if not project_id:
        raise RuntimeError("GOOGLE_CLOUD_PROJECT is not set.")

    model = ChatGoogleGenerativeAI(
        model=MODEL_ID,
        project=project_id,
        location="global",
        vertexai=True,
        max_retries=2,
    )
    return model.with_structured_output(
        schema=ANALYSIS_SCHEMA,
        method="json_schema",
    )

What does the state remember?

  • The state carries the original incident and its stable identifier through every node.
  • Policy fields record whether approval is required and why the rule matched.
  • The audit log records each completed decision in order.
  • Add the analysis node below build_analyst() by pasting this chunk:
def analyze_incident(state: IncidentState):
    analyst = build_analyst()
    analysis = analyst.invoke(
        [
            (
                "system",
                "You are an SRE incident analyst. Classify the incident and "
                "propose one reversible remediation step. Keep every field concise.",
            ),
            ("human", state["incident"]),
        ]
    )
    return {
        "analysis": analysis,
        "status": "analyzed",
        "audit_log": ["Model produced structured incident analysis."],
    }

What does the analysis node produce?

The node asks Gemini 3.7 Flash for structured incident analysis. It also starts the audit trail with a record that the model completed its assessment.

  • Add the deterministic policy node below analyze_incident() by pasting this chunk:
def policy_check(state: IncidentState):
    analysis = state["analysis"]
    searchable_text = (
        f"{state['incident']} {analysis['proposed_action']}"
    ).lower()
    severity_requires_review = analysis["severity"] in {"SEV1", "SEV2"}
    term_requires_review = any(
        term in searchable_text for term in RESTRICTED_TERMS
    )
    requires_approval = severity_requires_review or term_requires_review
    reason = (
        "High severity or restricted operational language detected."
        if requires_approval
        else "No high-risk policy rule matched."
    )
    return {
        "requires_approval": requires_approval,
        "policy_reason": reason,
        "status": "policy_checked",
        "audit_log": state["audit_log"] + [f"Policy check: {reason}"],
    }

Why use deterministic policy?

The model proposes an assessment. The policy owns the approval decision using fixed severity rules and restricted terms.

The raw incident is included in the search so terms such as production or restart still trigger review when the model paraphrases its proposal.

Seeing a syntax marker in the editor?

Check that each closing bracket from ANALYSIS_SCHEMA is present. Also confirm that the policy function sits below the completed analysis function.

Ask for targeted help with my app.py shows a syntax marker after adding the schema and policy functions.

Add durable interruption

A human-in-the-loop interrupt pauses graph execution and exposes a review payload. A thread-scoped SQLite checkpoint preserves the state so another Python process can resume it later.

  • Add the policy route and approval gate below policy_check() by pasting this chunk:
def policy_route(state: IncidentState) -> Literal["approval_gate", "publish_runbook"]:
    return "approval_gate" if state["requires_approval"] else "publish_runbook"


def approval_gate(state: IncidentState):
    decision = interrupt(
        {
            "question": "Approve this remediation proposal?",
            "incident_id": state["incident_id"],
            "incident": state["incident"],
            "analysis": state["analysis"],
            "policy_reason": state["policy_reason"],
            "allowed_decisions": ["approve", "reject"],
        }
    )
    approved = bool(decision)
    return {
        "decision": approved,
        "status": "approved" if approved else "rejected",
        "audit_log": state["audit_log"]
        + ["Human approved the proposal." if approved else "Human rejected the proposal."],
    }

How does the interrupt protect publishing?

  • The policy route sends high-risk incidents to approval_gate.
  • The interrupt exposes only JSON-serializable incident and policy data for review.
  • Execution pauses before the node can return a decision or reach publication.
  • Add the approval route and safe incident identifier helper below approval_gate() by pasting this chunk:
def approval_route(state: IncidentState) -> Literal["publish_runbook", "reject_runbook"]:
    return "publish_runbook" if state["decision"] else "reject_runbook"


def safe_incident_id(incident_id: str) -> str:
    safe_id = "".join(
        character
        for character in incident_id
        if character.isalnum() or character in {"-", "_"}
    )
    if not safe_id:
        raise ValueError("Incident ID must contain a letter or number.")
    return safe_id

Why sanitize the incident ID?

The thread identifier also becomes part of the runbook filename. The helper keeps letters, numbers, hyphens, and underscores so the side effect always targets a predictable path.

  • Start the publishing node below safe_incident_id() by pasting this preparation chunk:
def publish_runbook(state: IncidentState):
    RUNBOOK_DIR.mkdir(exist_ok=True)
    path = RUNBOOK_DIR / f"{safe_incident_id(state['incident_id'])}.md"
    analysis = state["analysis"]
    decision = "Approved" if state.get("decision") else "Not required by policy"
    audit_lines = "\n".join(f"- {entry}" for entry in state["audit_log"])

How does the runbook path stay stable?

The node derives one Markdown filename from the incident ID. Reusing that identifier targets the same path instead of creating duplicate artifacts.

  • Continue publish_runbook() by adding the file-writing section directly below audit_lines:
    path.write_text(
        "# Approved Incident Runbook\n\n"
        f"Incident ID: {state['incident_id']}\n\n"
        f"Incident: {state['incident']}\n\n"
        f"Severity: {analysis['severity']}\n\n"
        f"Summary: {analysis['summary']}\n\n"
        f"Proposed action: {analysis['proposed_action']}\n\n"
        f"Model risk: {analysis['risk']}\n\n"
        f"Policy result: {state['policy_reason']}\n\n"
        f"Human decision: {decision}\n\n"
        "Audit log:\n"
        f"{audit_lines}\n",
        encoding="utf-8",
    )

What goes into the audit artifact?

The Markdown file combines the original incident with the model analysis. It also records the deterministic policy result, the human decision, and the ordered audit log.

  • Finish publish_runbook() and add the rejection node by pasting this chunk:
    return {
        "status": "published",
        "runbook_path": str(path),
        "audit_log": state["audit_log"] + [f"Published runbook to {path}."],
    }


def reject_runbook(state: IncidentState):
    return {
        "status": "rejected_without_publish",
        "audit_log": state["audit_log"] + ["Publishing was skipped."],
    }

How do the two outcomes differ?

The publishing node writes the runbook and records its path. The rejection node reaches a terminal status without writing a file.

  • Connect the protected nodes through conditional edges by adding this graph builder:
def build_graph(checkpointer: SqliteSaver):
    builder = StateGraph(IncidentState)
    builder.add_node("analyze_incident", analyze_incident)
    builder.add_node("policy_check", policy_check)
    builder.add_node("approval_gate", approval_gate)
    builder.add_node("publish_runbook", publish_runbook)
    builder.add_node("reject_runbook", reject_runbook)

    builder.add_edge(START, "analyze_incident")
    builder.add_edge("analyze_incident", "policy_check")
    builder.add_conditional_edges("policy_check", policy_route)
    builder.add_conditional_edges("approval_gate", approval_route)
    builder.add_edge("publish_runbook", END)
    builder.add_edge("reject_runbook", END)

    return builder.compile(checkpointer=checkpointer)

What changes in the graph path?

Every incident now passes through policy_check after model analysis. Conditional edges either request approval or allow a low-risk proposal to continue.

Compiling with the checkpointer gives the interrupt a durable place to save thread state.

  • Add the event stream display helper below build_graph() by pasting this chunk:
def display_stream(stream):
    final_state = stream.output
    if stream.interrupted:
        print("WORKFLOW PAUSED")
        print(json.dumps(stream.interrupts[0].value, indent=2))
        return

    print("WORKFLOW COMPLETE")
    print(json.dumps(final_state, indent=2))

How is a pause made visible?

The event stream reports whether the graph stopped at an interrupt. The helper prints the review payload when paused and the final state when complete.

  • Add the command-line parser below display_stream() by pasting this chunk:
def parse_args():
    parser = argparse.ArgumentParser(
        description="Run a resumable incident-triage workflow."
    )
    subparsers = parser.add_subparsers(dest="command", required=True)

    start_parser = subparsers.add_parser("start")
    start_parser.add_argument("--thread", required=True)
    start_parser.add_argument("--incident", required=True)

    resume_parser = subparsers.add_parser("resume")
    resume_parser.add_argument("--thread", required=True)
    resume_parser.add_argument(
        "--decision",
        choices=["approve", "reject"],
        required=True,
    )

    return parser.parse_args()

What does the command interface capture?

The start command requires a thread ID and raw incident. The resume command requires the same thread ID plus an approval or rejection decision.

  • Add the first part of main() below parse_args() by pasting this chunk:
def main():
    args = parse_args()
    config = {"configurable": {"thread_id": args.thread}}

    with SqliteSaver.from_conn_string(DB_PATH) as checkpointer:
        graph = build_graph(checkpointer)

        if args.command == "start":
            graph_input = {
                "incident_id": args.thread,
                "incident": args.incident,
                "status": "new",
                "audit_log": [],
            }
        else:
            graph_input = Command(resume=args.decision == "approve")

How does thread-scoped recovery work?

The thread ID becomes part of the graph configuration. SqliteSaver uses that value to separate one incident's checkpoints from every other incident.

A new start creates initial state. A resume converts the operator's decision into the value returned by the paused interrupt.

  • Finish main() and add the script entry point by pasting this final chunk:
        stream = graph.stream_events(
            graph_input,
            config=config,
            durability="sync",
            version="v3",
        )
        display_stream(stream)


if __name__ == "__main__":
    main()

Why use synchronous durability?

The sync durability mode writes each checkpoint before the graph continues. The paused state is therefore available after the first Python process exits.

  • Save app.py in the Cloud Shell editor.

Having trouble completing app.py?

Check that every chunk appears in the same order as the full file below. Pay particular attention to indentation inside publish_runbook() and main().

Ask for a focused comparison with help me find indentation or ordering differences in my protected LangGraph app.py.

✔️ Awesome, I've got everything!

Your protected workflow code is ready. Make sure app.py is saved before starting the high-risk thread.

ⓧ I'd like to double check the full code

import argparse
import json
import os
from pathlib import Path
from typing import Literal, TypedDict

os.environ.setdefault("LANGGRAPH_STRICT_MSGPACK", "true")

from langchain_google_genai import ChatGoogleGenerativeAI
from langgraph.checkpoint.sqlite import SqliteSaver
from langgraph.graph import END, START, StateGraph
from langgraph.types import Command, interrupt

MODEL_ID = "gemini-3.7-flash"
DB_PATH = "incident_checkpoints.db"
RUNBOOK_DIR = Path("runbooks")
RESTRICTED_TERMS = (
    "delete",
    "disable",
    "drop",
    "production",
    "restart",
    "rotate",
)

ANALYSIS_SCHEMA = {
    "title": "IncidentAnalysis",
    "type": "object",
    "properties": {
        "severity": {
            "type": "string",
            "enum": ["SEV1", "SEV2", "SEV3", "SEV4"],
        },
        "summary": {"type": "string"},
        "proposed_action": {"type": "string"},
        "risk": {
            "type": "string",
            "enum": ["low", "high"],
        },
        "rationale": {"type": "string"},
    },
    "required": [
        "severity",
        "summary",
        "proposed_action",
        "risk",
        "rationale",
    ],
}


class IncidentState(TypedDict, total=False):
    incident_id: str
    incident: str
    analysis: dict[str, str]
    requires_approval: bool
    policy_reason: str
    decision: bool
    status: str
    audit_log: list[str]
    runbook_path: str


def build_analyst():
    project_id = os.environ.get("GOOGLE_CLOUD_PROJECT")
    if not project_id:
        raise RuntimeError("GOOGLE_CLOUD_PROJECT is not set.")

    model = ChatGoogleGenerativeAI(
        model=MODEL_ID,
        project=project_id,
        location="global",
        vertexai=True,
        max_retries=2,
    )
    return model.with_structured_output(
        schema=ANALYSIS_SCHEMA,
        method="json_schema",
    )


def analyze_incident(state: IncidentState):
    analyst = build_analyst()
    analysis = analyst.invoke(
        [
            (
                "system",
                "You are an SRE incident analyst. Classify the incident and "
                "propose one reversible remediation step. Keep every field concise.",
            ),
            ("human", state["incident"]),
        ]
    )
    return {
        "analysis": analysis,
        "status": "analyzed",
        "audit_log": ["Model produced structured incident analysis."],
    }


def policy_check(state: IncidentState):
    analysis = state["analysis"]
    searchable_text = (
        f"{state['incident']} {analysis['proposed_action']}"
    ).lower()
    severity_requires_review = analysis["severity"] in {"SEV1", "SEV2"}
    term_requires_review = any(
        term in searchable_text for term in RESTRICTED_TERMS
    )
    requires_approval = severity_requires_review or term_requires_review
    reason = (
        "High severity or restricted operational language detected."
        if requires_approval
        else "No high-risk policy rule matched."
    )
    return {
        "requires_approval": requires_approval,
        "policy_reason": reason,
        "status": "policy_checked",
        "audit_log": state["audit_log"] + [f"Policy check: {reason}"],
    }


def policy_route(state: IncidentState) -> Literal["approval_gate", "publish_runbook"]:
    return "approval_gate" if state["requires_approval"] else "publish_runbook"


def approval_gate(state: IncidentState):
    decision = interrupt(
        {
            "question": "Approve this remediation proposal?",
            "incident_id": state["incident_id"],
            "incident": state["incident"],
            "analysis": state["analysis"],
            "policy_reason": state["policy_reason"],
            "allowed_decisions": ["approve", "reject"],
        }
    )
    approved = bool(decision)
    return {
        "decision": approved,
        "status": "approved" if approved else "rejected",
        "audit_log": state["audit_log"]
        + ["Human approved the proposal." if approved else "Human rejected the proposal."],
    }


def approval_route(state: IncidentState) -> Literal["publish_runbook", "reject_runbook"]:
    return "publish_runbook" if state["decision"] else "reject_runbook"


def safe_incident_id(incident_id: str) -> str:
    safe_id = "".join(
        character
        for character in incident_id
        if character.isalnum() or character in {"-", "_"}
    )
    if not safe_id:
        raise ValueError("Incident ID must contain a letter or number.")
    return safe_id


def publish_runbook(state: IncidentState):
    RUNBOOK_DIR.mkdir(exist_ok=True)
    path = RUNBOOK_DIR / f"{safe_incident_id(state['incident_id'])}.md"
    analysis = state["analysis"]
    decision = "Approved" if state.get("decision") else "Not required by policy"
    audit_lines = "\n".join(f"- {entry}" for entry in state["audit_log"])

    path.write_text(
        "# Approved Incident Runbook\n\n"
        f"Incident ID: {state['incident_id']}\n\n"
        f"Incident: {state['incident']}\n\n"
        f"Severity: {analysis['severity']}\n\n"
        f"Summary: {analysis['summary']}\n\n"
        f"Proposed action: {analysis['proposed_action']}\n\n"
        f"Model risk: {analysis['risk']}\n\n"
        f"Policy result: {state['policy_reason']}\n\n"
        f"Human decision: {decision}\n\n"
        "Audit log:\n"
        f"{audit_lines}\n",
        encoding="utf-8",
    )
    return {
        "status": "published",
        "runbook_path": str(path),
        "audit_log": state["audit_log"] + [f"Published runbook to {path}."],
    }


def reject_runbook(state: IncidentState):
    return {
        "status": "rejected_without_publish",
        "audit_log": state["audit_log"] + ["Publishing was skipped."],
    }


def build_graph(checkpointer: SqliteSaver):
    builder = StateGraph(IncidentState)
    builder.add_node("analyze_incident", analyze_incident)
    builder.add_node("policy_check", policy_check)
    builder.add_node("approval_gate", approval_gate)
    builder.add_node("publish_runbook", publish_runbook)
    builder.add_node("reject_runbook", reject_runbook)

    builder.add_edge(START, "analyze_incident")
    builder.add_edge("analyze_incident", "policy_check")
    builder.add_conditional_edges("policy_check", policy_route)
    builder.add_conditional_edges("approval_gate", approval_route)
    builder.add_edge("publish_runbook", END)
    builder.add_edge("reject_runbook", END)

    return builder.compile(checkpointer=checkpointer)


def display_stream(stream):
    final_state = stream.output
    if stream.interrupted:
        print("WORKFLOW PAUSED")
        print(json.dumps(stream.interrupts[0].value, indent=2))
        return

    print("WORKFLOW COMPLETE")
    print(json.dumps(final_state, indent=2))


def parse_args():
    parser = argparse.ArgumentParser(
        description="Run a resumable incident-triage workflow."
    )
    subparsers = parser.add_subparsers(dest="command", required=True)

    start_parser = subparsers.add_parser("start")
    start_parser.add_argument("--thread", required=True)
    start_parser.add_argument("--incident", required=True)

    resume_parser = subparsers.add_parser("resume")
    resume_parser.add_argument("--thread", required=True)
    resume_parser.add_argument(
        "--decision",
        choices=["approve", "reject"],
        required=True,
    )

    return parser.parse_args()


def main():
    args = parse_args()
    config = {"configurable": {"thread_id": args.thread}}

    with SqliteSaver.from_conn_string(DB_PATH) as checkpointer:
        graph = build_graph(checkpointer)

        if args.command == "start":
            graph_input = {
                "incident_id": args.thread,
                "incident": args.incident,
                "status": "new",
                "audit_log": [],
            }
        else:
            graph_input = Command(resume=args.decision == "approve")

        stream = graph.stream_events(
            graph_input,
            config=config,
            durability="sync",
            version="v3",
        )
        display_stream(stream)


if __name__ == "__main__":
    main()

How to compare the file

Compare the imports, constants, functions, and entry point in order. Your saved app.py should match this reference exactly.

Prove publishing is blocked

The safety boundary matters only if the runbook stays absent while the workflow waits. This check starts a fresh high-risk thread and tests the file path immediately after the process pauses.

The reset removes only this project's checkpoint database and generated runbook directory. Your source files and virtual environment remain intact.

Before you run this, consider whether the restricted production language sends the thread to publication or approval.

  • Start a clean high-risk thread and test the protected runbook path by running these commands:
rm -f incident_checkpoints.db
rm -rf runbooks
python3 app.py start --thread incident-001 --incident "Production checkout is failing for all users. The proposed emergency action is to restart the production database."
test ! -f runbooks/incident-001.md && echo "No runbook published before approval"

What do these commands prove?

  • The first two commands remove old checkpoints and runbooks so the test begins from a known state.
  • The start command analyzes the incident and applies deterministic policy to thread incident-001.
  • The final command prints a safety confirmation only when runbooks/incident-001.md does not exist.

You should see WORKFLOW PAUSED followed by an approval payload. You should then see No runbook published before approval.

That is the control working. Your high-risk thread is safely checkpointed without publishing an unreviewed plan.

Did the workflow fail to pause?

Confirm that the incident includes the restricted terms production and restart. Also check that policy_route() sends a true requires_approval value to approval_gate.

Ask for help with my protected incident workflow did not pause before publishing.

Your approval gate now survives the end of the Python process. Next, you will resume incident-001 from its saved checkpoint and publish the approved audit artifact.

Resume the Approved Workflow

In your Google Cloud Shell session from earlier, your LangGraph workflow holds incident-001 at approval_gate. The protected path has stopped before publishing.

A useful approval gate must survive process boundaries. A new Python process can recover the matching SQLite checkpoint through the thread ID.

In this step, get ready to:
  • Resume the saved incident from a separate Python process.
  • Inspect the approved Markdown runbook.
  • Prove that a rejected thread finishes without publishing.
Resume the approved thread

The saved thread contains every value produced before the pause. Its pending node is the approval gate.

Before you run this, predict whether the new Python process will reuse the saved analysis.

  • Resume incident-001 with an approval decision by running this command:
python3 app.py resume --thread incident-001 --decision approve

What Does This Resume Prove?

  • The matching thread_id selects the saved checkpoint for incident-001.
  • The approval choice becomes the Boolean result used by the paused node.
  • approval_route sends the workflow to publish_runbook.
  • The terminal prints WORKFLOW COMPLETE with the status published.
  • Confirm the completed state contains runbooks/incident-001.md as the runbook path.

Workflow Pauses Again or Cannot Resume?

Check that the command uses the same incident-001 thread ID that paused earlier. A different ID cannot select that checkpoint.

Run the command from ~/langgraph-incident-triage with the existing virtual environment active.

Help me diagnose why my saved thread will not resume.

You have crossed the process boundary. The saved incident continued without another model-analysis request.

Inspect the approved audit trail

The Markdown runbook turns the completed workflow state into a reviewable incident artifact. Its stable incident filename also makes the publishing side effect idempotent.

Before you inspect the file, predict which details prove that policy review happened before publishing.

  • Print the approved runbook in your terminal by running this command:
cat runbooks/incident-001.md

What Should the Artifact Prove?

  • Incident preserves the original production report.
  • Severity records the model classification.
  • Summary records the concise incident assessment.
  • Proposed action records the suggested remediation.
  • Model risk records the structured risk value.
  • Policy result records the deterministic reason for review.
  • Human decision: Approved proves that publishing followed approval.
  • Audit log preserves the model step before the policy step. The human decision follows both.
  • Confirm the completed JSON from the resume command contains the model's rationale field.
  • Confirm the runbook contains the approved human decision.
  • Confirm the audit entries appear in workflow order.

Runbook Missing or Incomplete?

Confirm the resume output reported the published status. The file is created only after the publish node runs.

Check that you are still inside ~/langgraph-incident-triage. The relative runbooks/ path starts from that directory.

Help me find my approved incident runbook.

Approval is meaningful because rejection follows its own terminal path. A separate thread lets you test that route without changing the approved artifact.

  • Start a separate high-risk thread by running this command:
python3 app.py start --thread incident-rejected-001 --incident "Production checkout is failing for all users. The proposed emergency action is to restart the production database."

Why Use a Separate Thread?

The new thread receives its own checkpoint history. The command pauses at the approval gate without changing incident-001.

  • Confirm the terminal prints WORKFLOW PAUSED for incident-rejected-001.

Before you resume this thread, predict whether rejection will create a Markdown runbook.

  • Resume the separate thread with a rejection decision by running this command:
python3 app.py resume --thread incident-rejected-001 --decision reject

What Does Rejection Change?

  • The rejection choice becomes a false Boolean value in approval_gate.
  • approval_route sends the saved thread to reject_runbook.
  • The final state reports rejected_without_publish.
  • The audit log records Publishing was skipped..
  • Inspect the runbooks/ directory with the file browser in your existing Cloud Shell session.
  • Confirm the directory contains incident-001.md.
  • Confirm the directory does not contain incident-rejected-001.md.

Rejected Thread Created a Runbook?

Check that the resume command used --decision reject for incident-rejected-001.

Make sure you are checking the rejected thread's filename. The approved incident-001.md file should remain in the same directory.

Help me trace why my rejected thread reached the publish node.

That closes the approval loop. Your workflow now demonstrates durable approval, controlled publishing, and a terminal rejection path.

Secret mission

Inspect a Saved Checkpoint

A durable approval gate needs a safe way to reveal why a workflow is waiting. Build a checkpoint inspector that reports saved state without resuming the thread or publishing a runbook.

Clean Up Your Resources

Clean Up Your Resources

Your saved files in Google Cloud Shell create no ongoing compute cost while the workflow is idle. Gemini requests remain usage-based, so decide whether to keep the project, pause your work, or delete its local resources.

Resources you used:

  • Cloud Shell project directory at ~/langgraph-incident-triage containing requirements.txt, naive_app.py, app.py, and inspect_state.py.
  • Python virtual environment in ~/langgraph-incident-triage/.venv.
  • SQLite checkpoint database at ~/langgraph-incident-triage/incident_checkpoints.db containing the completed incident-001 thread plus the paused inspection-001 thread.
  • Generated Markdown runbooks under ~/langgraph-incident-triage/runbooks/.

Keep everything running

Retain the project if you want to demonstrate the approval workflow or inspect its saved threads again.

  • Leave ~/langgraph-incident-triage in Cloud Shell's persistent home storage.
  • Keep incident_checkpoints.db to preserve the completed and paused workflow states.
  • Leave the Vertex AI API enabled for the earlier gateway project.
  • Run the workflow again only when you want to make another Gemini request.

Pause - I'll come back to this later

Pausing requires no shutdown because the project has no continuously running service. The saved inspection-001 checkpoint remains waiting inside the database without executing code.

  • Leave ~/langgraph-incident-triage intact in Cloud Shell.
  • Avoid starting or resuming workflows until you return.
  • Keep the Vertex AI API enabled if the earlier gateway project still uses it.

Delete - I don't want to use this again

Deleting the directory is final. The command removes this project's files without deleting your existing Google Cloud project.

  • Remove all local project resources by running this command:
rm -rf ~/langgraph-incident-triage

What Does This Cleanup Command Do?

The rm -rf command recursively removes the project directory. Its target includes the virtual environment, scripts, checkpoint database, and generated runbooks.

  • Confirm ~/langgraph-incident-triage no longer appears in the Cloud Shell file browser.

Cleanup complete. Your workflow files and saved checkpoint history are now removed from Cloud Shell.

The Vertex AI API remains enabled because the earlier gateway project may still use it.

Still See the Project Directory?

  • Check that the cleanup command targeted ~/langgraph-incident-triage exactly.
  • Help me troubleshoot this cleanup command.

Nice Work!

Nice Work!

You made it. Your LangGraph incident-triage workflow now blocks high-risk runbooks until an operator approves the proposed action.

You've learned how to:

  • Expose the danger of an unsafe auto-publisher by running a naive graph that creates naive_runbook.md before human review.
  • Protect high-risk incidents with an approval boundary. Structured Gemini analysis feeds deterministic policy checks. A human-in-the-loop interrupt pauses publishing.
  • Resume the exact saved thread from a SQLite checkpoint in a new process. The approved path produces an auditable Markdown runbook under runbooks/.
  • Secret Mission: Inspect the saved inspection-001 checkpoint through inspect_state.py. The inspector reveals the saved state without resuming the workflow. It leaves runbooks/inspection-001.md absent.

Ready to quiz yourself?