Build an Agentic RAG Copilot

Build an evidence-grounded incident-response API with Gemini and LangGraph.

Introduction

30 Second Summary

During an outage, the loudest alert can pull attention away from the evidence that explains what is failing. A rushed guess can waste precious minutes when the team needs a clear next action.

In this project, you will build a local incident-response copilot that answers operational questions from synthetic runbooks through cloud AI. The Gemini API handles embeddings plus generated responses while the service stays inspectable on your Mac.

What You'll Build

Your finished demo takes a compound incident question from FastAPI's interactive page before returning an evidence-backed action plan with source identifiers plus its full decision route.

By the end of this project, you'll have:

  • A local incident-response API for operational questions. Every response includes an answer grounded in synthetic runbooks.
  • An inspectable corrective Agentic RAG workflow in LangGraph. You can show its planner route, retrieved sources, review result, plus bounded retry.
  • A machine-readable evaluation report that measures expected-source retrieval, citation presence, route behavior, plus observed cloud latency.
  • Secret Mission: Add a queue-backlog runbook to the knowledge base. Prove the new incident domain with a regression case.

Are there any prerequisites?

You need a Google account plus internet access for cloud model calls. Your Mac also needs Python 3.10 or newer plus Visual Studio Code.

Before We Start

Before the hands-on work begins, this step gives you a moment to define the incident-response copilot you are building. You will connect its evidence-backed synthetic runbooks to safer guidance for engineers handling operational incidents.

Set Up the Lightweight Cloud AI Stack

Your Mac cannot run Ollama comfortably. Cloud inference lets the copilot use the Gemini API without downloading a local model.

The application stays on your Mac. Your synthetic data also remains easy to inspect locally.

In this step, get ready to:
  • Confirm that Python 3.10 or newer is available alongside Visual Studio Code.
  • Create a Gemini API key through Google AI Studio without enabling paid billing.
  • Create a local virtual environment with the pinned dependencies.
  • Verify that the cloud model returns READY.
Confirm Python and Visual Studio Code

The pinned libraries require Python 3.10 or newer. Checking the runtime first prevents dependency errors during installation.

  • Press Cmd+Space on macOS or the Windows key on Windows to open system search.
  • Type Terminal into the search field.
  • Press Enter to open Terminal.
  • Check your installed Python version by running this command:
python --version

What does this command do?

The command asks your shell which Python version runs when you use python. The result determines whether the pinned packages support your runtime.

Match the version shown in your terminal to the relevant path below.

✔️ I see version 3.10 or higher

Your Python runtime supports every pinned package in this project. That compatibility check is complete.

ⓧ I see an older version

The installed runtime is below the minimum supported version. The current macOS installer from Python.org provides a compatible release.

  • Download the current installer from the Python macOS releases page.
  • Complete the installer using its default options.
  • Close Terminal after the installation finishes.
  • Press Cmd+Space to reopen system search.
  • Type Terminal into the search field.
  • Press Enter to open a fresh Terminal session.
  • Confirm that the upgraded runtime is active by running:
python --version

Why check again?

A fresh Terminal session loads the shell configuration created by the installer. The command confirms which Python executable your shell now resolves.

Continue when the terminal reports Python 3.10 or newer.

Still seeing the older version?

Your shell may still resolve an older Python installation first. Follow the Python.org macOS setup guidance until python reports a compatible version.

Help me update the Python version used by my macOS shell.

ⓧ Command not found

Your shell cannot find a Python installation under the required command. Install the current macOS release before continuing.

  • Download the current installer from the Python macOS releases page.
  • Complete the installer using its default options.
  • Close Terminal after the installation finishes.
  • Press Cmd+Space to reopen system search.
  • Type Terminal into the search field.
  • Press Enter to open a fresh Terminal session.
  • Confirm that Python is available by running:
python --version

What confirms the installation?

A version number confirms that the shell can now find Python. Continue when that version is 3.10 or newer.

Python still unavailable?

Restart Terminal after completing the installer. A fresh shell can load the path created during installation.

Help me make Python available in my macOS Terminal.

The code and dependency files need an editor. Confirm that Visual Studio Code is ready before creating the project.

  • Press Cmd+Space on macOS or the Windows key on Windows to open system search.
  • Type Visual Studio Code into the search field.
  • Press Enter to open Visual Studio Code.

You should see the Visual Studio Code window. Your editor and compatible Python runtime are ready.

Visual Studio Code missing?

Install the macOS build from the official Visual Studio Code download page. Reopen it through system search after installation.

Help me install Visual Studio Code on my Mac.

Create a Gemini API key

An API key authorizes the local application to call the cloud models. You will keep this credential in one terminal session.

Paid billing stays disabled for this project. The main path uses the Gemini API Free Tier.

  • Open Google AI Studio in your browser.
  • Sign in with your Google account.
  • Accept the service terms if Google AI Studio displays them.
  • Select Dashboard.
  • Select API Keys.
  • Create an API key for the available Google Cloud project.
  • Leave paid billing disabled.

Why use synthetic runbooks?

Gemini Free Tier content may be used to improve Google's products. This project sends only synthetic runbooks and practice incident questions.

Keep confidential operational data outside this learning environment. Synthetic evidence gives you the full workflow without exposing real incidents.

Your key grants access to the Gemini API. Keep it out of project files and source control.

  • Copy the new API key from Google AI Studio.
  • Switch back to the Terminal session from earlier.
  • Set the key for this terminal session by replacing the plain-text placeholder in the command below before you run it:
export GEMINI_API_KEY="your-api-key-here"

What does this command do?

The command creates the GEMINI_API_KEY environment variable in the current Terminal session. The Gemini client libraries detect this variable when they make a request.

Closing this Terminal session removes the exported value. No credential is written into your project files.

Key setup feels unclear?

Make sure you replaced the placeholder with the complete key. Keep the quotation marks around the value.

Help me set my Gemini API key safely in the current Terminal session.

Create and verify the local environment

The local environment isolates this project's pinned packages from other Python projects. The original Terminal session also retains your API key.

  • Switch back to the Terminal session from earlier.
  • Move to your Desktop by running this command:
cd ~/Desktop

What does this command do?

The cd command changes the current location to your Desktop. The new project folder will be easy to find there.

  • Create the cloud-agentic-rag-copilot folder and confirm its location by running:
mkdir cloud-agentic-rag-copilot
cd cloud-agentic-rag-copilot
pwd

What do these commands do?

  • The first command creates the project folder on your Desktop.
  • The second command moves the Terminal session into that folder.
  • The final command prints the current folder path for confirmation.

The printed path should end with /Desktop/cloud-agentic-rag-copilot.

Folder creation failed?

A folder with the same name may already exist on your Desktop. Remove the unused duplicate through Finder or continue inside the existing empty folder.

Help me fix the project folder setup on my Desktop.

  • Switch back to Visual Studio Code from earlier.
  • Click File in the top menu bar.
  • Click Open Folder.
  • Select the cloud-agentic-rag-copilot folder on your Desktop.
  • Click Open in the folder picker.

The Explorer sidebar should now show cloud-agentic-rag-copilot as the open folder.

  • Select the New File control in the Explorer sidebar.
  • Enter requirements.txt as the file name.
  • Add the pinned dependencies by pasting this exact content into requirements.txt:
fastapi==0.142.2
google-genai==2.28.0
langchain-core==1.6.6
langchain-google-genai==4.4.0
langgraph==1.2.13
pydantic==2.13.5
uvicorn[standard]==0.54.0

What does this file control?

Each line pins one dependency to an exact version. Pinning makes your installation reproducible across future runs.

The list covers the API server, Gemini clients, vector-store interfaces, graph orchestration, data validation, and development server.

  • Save requirements.txt by pressing Cmd+S on macOS or Ctrl+S on Windows.
  • Confirm that requirements.txt remains listed beneath cloud-agentic-rag-copilot in the Explorer sidebar.

Dependency list looks different?

Check each package name and version against the reference below. A missing character can stop the installation.

Help me compare my requirements file with the expected dependency list.

✔️ Awesome, I've got everything!

Your dependency file is saved with every required package pin.

ⓧ I'd like to double check the full code

fastapi==0.142.2
google-genai==2.28.0
langchain-core==1.6.6
langchain-google-genai==4.4.0
langgraph==1.2.13
pydantic==2.13.5
uvicorn[standard]==0.54.0

How should this file look?

This reference is the complete requirements.txt file for the project. Every package name and version should match exactly.

Installing the pinned packages can take a few minutes while Python downloads each dependency. A quiet pause during the installation is normal.

  • Switch back to the original Terminal session.
  • Create the virtual environment, activate it, and install the pinned dependencies by running:
python -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt

What do these commands do?

  • The first command creates an isolated Python environment in the .venv folder.
  • The second command activates that environment in the current Terminal session.
  • The final command installs every package pinned in requirements.txt.

Your Terminal prompt should begin with (.venv) after activation. The installation should complete without a dependency error.

Dependency installation failed?

Confirm that the terminal prompt begins with (.venv). Check that your Python version meets the minimum requirement.

A network interruption can also stop package downloads. Run the installation again after confirming your internet connection.

Help me diagnose my virtual environment or dependency installation.

Before you run this check, ask yourself whether the cloud model can answer from the new environment.

  • Verify the model connection by running this command in the active virtual environment:
python -c 'from langchain_google_genai import ChatGoogleGenerativeAI; print(ChatGoogleGenerativeAI(model="gemini-3.8-flash", temperature=1.0).invoke("Reply only READY").text)'

What does this command prove?

The command imports ChatGoogleGenerativeAI from the installed integration. It creates a client for gemini-3.8-flash.

The invocation sends a small request through the Gemini API. Its response proves that the package installation and current-session API key work together.

You should see READY in the terminal. That is the hard part complete: your lightweight cloud AI stack can now take model requests.

Model check did not return READY?

Confirm that you ran the command in the same Terminal session where you exported GEMINI_API_KEY. Confirm that (.venv) appears at the start of the prompt.

Check your internet connection if the request cannot reach the provider. Google AI Studio also shows the active limits for your project.

Help me debug the Gemini model verification command.

Your local environment can now reach Gemini without downloading a model. Next up, you will turn synthetic runbooks into searchable evidence.

Build Semantic Runbook Search

Your lightweight Gemini API setup already proves that this Mac can call a cloud model without downloading one.

Incident descriptions often use different words for the same operational problem. In this step, embeddings connect those descriptions to the relevant synthetic runbook.

In this step, get ready to:
  • Store three synthetic incident runbooks with stable source identifiers.
  • Convert each runbook into a searchable cloud embedding.
  • Return the source identifier for the top semantic match.
Add synthetic incident runbooks

A runbook captures symptoms plus the actions used to resolve them. Stable source identifiers let later answers point back to the evidence they used.

  • In the VS Code file sidebar, create a folder named data inside your project folder.
  • Inside data, create a file named runbooks.json.
  • Add the three synthetic runbooks by pasting this code into data/runbooks.json:
[
  {
    "source": "api-503",
    "title": "API 503 Response Runbook",
    "content": "Symptoms: clients receive HTTP 503 responses and availability alerts fire. Actions: check service health, compare the start time with recent deployments, inspect upstream dependency health, and either add healthy capacity or roll back the suspected release. Escalate when errors continue after capacity and rollback checks."
  },
  {
    "source": "db-pool",
    "title": "Database Connection Pool Exhaustion Runbook",
    "content": "Symptoms: requests wait for database connections, pool timeout logs increase, and active connections remain near the configured maximum. Actions: identify long-running queries, verify connections are returned to the pool, reduce request pressure, and increase the pool only after checking database capacity. Escalate when blocked transactions or database saturation continue."
  },
  {
    "source": "disk-capacity",
    "title": "Disk Capacity Alert Runbook",
    "content": "Symptoms: filesystem usage exceeds the operational threshold and write failures may begin. Actions: locate rapidly growing logs or temporary files, rotate or archive safe candidates, verify retention settings, and expand the volume when cleanup cannot restore headroom. Never delete unknown database files."
  }
]

How is each runbook structured?

  • The source value gives each runbook a stable identifier that can appear in citations.
  • The title value names the incident procedure.
  • The content value combines observable symptoms with response actions.
  • Save data/runbooks.json.
  • Review the file for matching braces plus matching quotation marks.

You should see one top-level array containing the api-503, db-pool, and disk-capacity source identifiers.

Seeing JSON syntax markers?

Check that each runbook object ends with a comma except the final object. Confirm that every property value stays inside double quotation marks.

Ask for help checking the file without changing its runbook content: Help me find the JSON syntax problem in my synthetic runbook file.

✔️ Awesome, I've got everything!

Your three synthetic runbooks are saved with stable source identifiers.

ⓧ I'd like to double check the full code

[
  {
    "source": "api-503",
    "title": "API 503 Response Runbook",
    "content": "Symptoms: clients receive HTTP 503 responses and availability alerts fire. Actions: check service health, compare the start time with recent deployments, inspect upstream dependency health, and either add healthy capacity or roll back the suspected release. Escalate when errors continue after capacity and rollback checks."
  },
  {
    "source": "db-pool",
    "title": "Database Connection Pool Exhaustion Runbook",
    "content": "Symptoms: requests wait for database connections, pool timeout logs increase, and active connections remain near the configured maximum. Actions: identify long-running queries, verify connections are returned to the pool, reduce request pressure, and increase the pool only after checking database capacity. Escalate when blocked transactions or database saturation continue."
  },
  {
    "source": "disk-capacity",
    "title": "Disk Capacity Alert Runbook",
    "content": "Symptoms: filesystem usage exceeds the operational threshold and write failures may begin. Actions: locate rapidly growing logs or temporary files, rotate or archive safe candidates, verify retention settings, and expand the volume when cleanup cannot restore headroom. Never delete unknown database files."
  }
]
Build the cloud embedding index

An embedding turns text into a numeric vector that preserves semantic similarity. Gemini Embedding 2 creates those vectors in the cloud.

A local vector store holds the returned vectors in memory. This keeps retrieval inspectable without adding a database.

  • In the VS Code file sidebar, create retrieval.py inside your project folder.
  • Add the imports plus fixed retrieval settings by pasting this code into retrieval.py:
import json
import re
from pathlib import Path

from google import genai
from google.genai import types
from langchain_core.embeddings import Embeddings
from langchain_core.vectorstores import InMemoryVectorStore

DATA_FILE = Path(__file__).parent / "data" / "runbooks.json"
SOURCE_PATTERN = re.compile(r"^\[source:([a-z0-9-]+)\]")
EMBEDDING_MODEL = "gemini-embedding-2"

What does this code prepare?

  • The imports provide JSON loading plus source-label parsing.
  • The DATA_FILE path points to the synthetic runbooks beside the script.
  • The SOURCE_PATTERN expression captures a stable source identifier from a retrieved document.
  • The EMBEDDING_MODEL value keeps the cloud model identifier in one place.
  • Save retrieval.py.
  • Review the top of retrieval.py for all three constant names.

You should see DATA_FILE, SOURCE_PATTERN, and EMBEDDING_MODEL directly below the imports.

Seeing unresolved imports?

Confirm that VS Code is using the Python interpreter from the active .venv environment. The pinned packages from the previous step provide these imports.

Get help checking the selected environment: Help me connect VS Code to my active virtual environment.

The embedding adapter gives LangChain one interface for document vectors plus query vectors. Each path adds a retrieval prefix before calling the cloud endpoint.

  • Below EMBEDDING_MODEL = "gemini-embedding-2", add the embedding adapter by pasting this code:
class GeminiEmbeddings(Embeddings):
    def __init__(self) -> None:
        self.client = genai.Client()

    def _embed(self, text: str) -> list[float]:
        result = self.client.models.embed_content(
            model=EMBEDDING_MODEL,
            contents=text,
            config=types.EmbedContentConfig(output_dimensionality=768),
        )
        return list(result.embeddings[0].values)

    def embed_documents(self, texts: list[str]) -> list[list[float]]:
        return [
            self._embed(f"title: incident runbook | text: {text}")
            for text in texts
        ]

    def embed_query(self, text: str) -> list[float]:
        return self._embed(f"task: search result | query: {text}")

How does the adapter work?

  • The Google Gen AI SDK client reads the current terminal session's configured API key.
  • The _embed() method requests a 768-dimensional vector.
  • The embed_documents() method prefixes each runbook as searchable reference material.
  • The embed_query() method prefixes the incident description as a search query.
  • Save retrieval.py.

Before you run this checkpoint, do you expect Python to need a local model file?

  • Check the adapter definitions by running this command in the active project terminal:
python retrieval.py

What does this check prove?

You should return to the terminal prompt without a traceback. Python can import the dependencies plus read every adapter definition without a local model runtime.

Does the script stop with an error?

Check the indentation inside GeminiEmbeddings. Confirm that each method stays aligned within the class.

Get help comparing the class structure: Help me check the indentation in my embedding adapter.

The adapter needs one consistently formatted document for each JSON object. The source label becomes part of the indexed text so later code can recover it.

  • Below the GeminiEmbeddings class, add the runbook loader by pasting this code:
def load_documents() -> list[str]:
    runbooks = json.loads(DATA_FILE.read_text(encoding="utf-8"))
    return [
        f"[source:{item['source']}]\nTitle: {item['title']}\n{item['content']}"
        for item in runbooks
    ]

What does the loader produce?

The loader reads the JSON array into Python objects. It converts each object into a document containing its source label plus title plus response content.

  • Save retrieval.py.
  • Review the returned document template for the source plus Title labels.

Each loaded document now starts with a machine-readable source label that the retrieval result preserves.

Does the loader look misaligned?

Keep the list comprehension inside the function's returned list. Confirm that the closing bracket aligns with return.

Ask for help checking the loader structure: Help me fix the load_documents function structure.

The in-memory index embeds every loaded runbook when Python imports the module. Restarting the process rebuilds this index from data/runbooks.json.

  • Below load_documents(), build the in-process index by pasting this code:
def build_vector_store() -> InMemoryVectorStore:
    vector_store = InMemoryVectorStore(GeminiEmbeddings())
    vector_store.add_texts(load_documents())
    return vector_store


VECTOR_STORE = build_vector_store()

How is the index built?

  • The InMemoryVectorStore uses your adapter to turn runbooks into vectors.
  • The add_texts() call indexes all three formatted documents.
  • The VECTOR_STORE constant keeps the finished index available to search functions.
  • Save retrieval.py.

The first cloud embedding call can pause briefly while the three runbook vectors are created.

Before you run this checkpoint, do you expect the index to use a downloaded model or the configured cloud endpoint?

  • Build the in-memory index by running this command:
python retrieval.py

What should happen?

You should return to the terminal prompt without a traceback after the embeddings finish. This confirms that Python can read the API key plus index all three runbooks.

Does indexing fail?

Confirm that you are using the same terminal session where the API key was set. Check that data/runbooks.json sits inside the project folder beside retrieval.py.

Get help isolating the failed boundary: Help me debug the runbook indexing failure.

Search the runbooks and verify the match

Cosine-similarity retrieval compares the query vector with each stored runbook vector. The closest document becomes the top semantic match even when the question uses different wording.

  • Below VECTOR_STORE = build_vector_store(), add the semantic search function by pasting this code:
def search_runbooks(question: str, k: int = 1) -> list[str]:
    documents = VECTOR_STORE.similarity_search(query=question, k=k)
    return [document.page_content for document in documents]

What does the search return?

The k value controls how many matching runbooks come back. Each result is converted from a stored document object into its original text.

  • Save retrieval.py.
  • Review search_runbooks() for the similarity_search(query=question, k=k) call.

Your search function now accepts natural incident language plus a result count.

Is the search method missing?

Check that VECTOR_STORE is capitalized consistently. Confirm that the search method is indented inside search_runbooks().

Ask for help comparing the method call: Help me fix the semantic search function.

The final helpers parse the stable source label from each retrieved document. The script entry point uses a database-pool symptom to make the semantic result visible.

  • Below search_runbooks(), add the source helpers plus runnable search by pasting this code:
def extract_source(document: str) -> str:
    match = SOURCE_PATTERN.search(document)
    return match.group(1) if match else "unknown"


def source_ids(documents: list[str]) -> list[str]:
    return [extract_source(document) for document in documents]


if __name__ == "__main__":
    sample_question = "Requests are waiting because the database pool has no free connections."
    matches = search_runbooks(sample_question, k=1)
    print(f"Source: {extract_source(matches[0])}")
    print(matches[0])

How does the result become inspectable?

  • The extract_source() function reads the source identifier at the start of a document.
  • The source_ids() function applies that parser to a list of retrieved documents.
  • The script entry point searches for one result plus prints its source identifier plus full runbook text.
  • Save retrieval.py.

Before you run the search, which runbook should be closest to a question about waiting for database connections?

  • Run the completed semantic search with this command:
python retrieval.py

What should you see?

The terminal should print Source: db-pool first. Below it, you should see the database connection-pool runbook with symptoms plus response actions.

That result proves the cloud embedding connected the sample wording to the relevant synthetic procedure.

Did a different source appear?

Confirm that all three runbook objects match the supplied content. Check that the query prefix remains inside embed_query().

Get help tracing the retrieved result: Help me debug why db-pool is not the top semantic match.

✔️ Awesome, I've got everything!

Your retrieval script embeds the synthetic runbooks plus returns the relevant database-pool source.

ⓧ I'd like to double check the full code

import json
import re
from pathlib import Path

from google import genai
from google.genai import types
from langchain_core.embeddings import Embeddings
from langchain_core.vectorstores import InMemoryVectorStore

DATA_FILE = Path(__file__).parent / "data" / "runbooks.json"
SOURCE_PATTERN = re.compile(r"^\[source:([a-z0-9-]+)\]")
EMBEDDING_MODEL = "gemini-embedding-2"


class GeminiEmbeddings(Embeddings):
    def __init__(self) -> None:
        self.client = genai.Client()

    def _embed(self, text: str) -> list[float]:
        result = self.client.models.embed_content(
            model=EMBEDDING_MODEL,
            contents=text,
            config=types.EmbedContentConfig(output_dimensionality=768),
        )
        return list(result.embeddings[0].values)

    def embed_documents(self, texts: list[str]) -> list[list[float]]:
        return [
            self._embed(f"title: incident runbook | text: {text}")
            for text in texts
        ]

    def embed_query(self, text: str) -> list[float]:
        return self._embed(f"task: search result | query: {text}")


def load_documents() -> list[str]:
    runbooks = json.loads(DATA_FILE.read_text(encoding="utf-8"))
    return [
        f"[source:{item['source']}]\nTitle: {item['title']}\n{item['content']}"
        for item in runbooks
    ]


def build_vector_store() -> InMemoryVectorStore:
    vector_store = InMemoryVectorStore(GeminiEmbeddings())
    vector_store.add_texts(load_documents())
    return vector_store


VECTOR_STORE = build_vector_store()


def search_runbooks(question: str, k: int = 1) -> list[str]:
    documents = VECTOR_STORE.similarity_search(query=question, k=k)
    return [document.page_content for document in documents]


def extract_source(document: str) -> str:
    match = SOURCE_PATTERN.search(document)
    return match.group(1) if match else "unknown"


def source_ids(documents: list[str]) -> list[str]:
    return [extract_source(document) for document in documents]


if __name__ == "__main__":
    sample_question = "Requests are waiting because the database pool has no free connections."
    matches = search_runbooks(sample_question, k=1)
    print(f"Source: {extract_source(matches[0])}")
    print(matches[0])

Your semantic search now finds operational evidence from varied incident language. Next, you'll connect one retrieved runbook to generation plus expose the limit of a top-one retrieval flow.

Expose the Single-Pass RAG Limitation

Your semantic retrieval layer now maps an incident description to its closest runbook. That proves meaning-based search works.

A compound incident needs evidence from multiple procedures. A single-pass RAG flow limited to one result can create an answer with a hidden evidence gap.

You will connect top-one retrieval to the Gemini API for answer generation. The compound question makes any missing evidence visible in the terminal.

In this step, get ready to:
  • Configure a top-one baseline for a compound incident question.
  • Run the baseline to reveal its missing evidence.
Configure the baseline question and model
  • Switch back to Visual Studio Code with the incident-response project from the previous step.
  • Create a baseline.py file beside retrieval.py using the file controls in the sidebar.
  • Add the model configuration plus compound question to baseline.py by copying this code:
from langchain_google_genai import ChatGoogleGenerativeAI

from retrieval import search_runbooks, source_ids

MODEL = ChatGoogleGenerativeAI(
    model="gemini-3.8-flash",
    temperature=1.0,
    max_retries=2,
)

COMPOUND_QUESTION = (
    "The API is returning 503 errors and the database connection pool is exhausted. "
    "What should I do?"
)

What does this code set up?

  • The imports connect the cloud chat model to the retrieval helpers from retrieval.py.
  • MODEL configures gemini-3.8-flash with temperature=1.0 plus max_retries=2.
  • COMPOUND_QUESTION deliberately combines an API availability symptom with database connection-pool exhaustion.
  • Save baseline.py.
  • Confirm baseline.py appears beside retrieval.py in the file sidebar.

Can't find baseline.py?

  • Confirm the file name is exactly baseline.py.
  • Move the file beside retrieval.py if it was created in another folder.

Help me locate my file

Build and test the single-pass flow

The baseline needs a complete retrieve-then-generate path. Its evidence boundary stays fixed at one retrieved document.

  • Place the answer flow directly below COMPOUND_QUESTION in baseline.py by copying this code:
def message_text(response: object) -> str:
    text = getattr(response, "text", None)
    return text if isinstance(text, str) else str(getattr(response, "content", response))


def baseline_answer(question: str) -> dict[str, object]:
    documents = search_runbooks(question, k=1)
    context = "\n\n".join(documents)
    response = MODEL.invoke(
        [
            (
                "system",
                "Answer only from the supplied runbook context. State when the context is incomplete.",
            ),
            ("human", f"Context:\n{context}\n\nQuestion:\n{question}"),
        ]
    )
    return {
        "question": question,
        "answer": message_text(response),
        "sources": source_ids(documents),
    }


if __name__ == "__main__":
    result = baseline_answer(COMPOUND_QUESTION)
    print(f"Question: {result['question']}")
    print(f"Sources: {result['sources']}")
    print(f"Answer:\n{result['answer']}")

How does the baseline work?

  • message_text() extracts a string from the model response.
  • baseline_answer() retrieves exactly one runbook with k=1.
  • The model receives only the retrieved document as incident evidence.
  • The returned dictionary preserves the question plus the answer plus the source list.
  • The direct-execution block runs the compound question when you execute baseline.py.
  • Save baseline.py.

Before you run this, do you expect one runbook to cover both symptoms equally well?

  • Run the baseline from your active terminal with this command:
python baseline.py

The Evidence Gap Is Visible

Your terminal prints the compound question. The Sources: line contains exactly one source ID.

The answer has runbook evidence for the symptom represented by that source. The other symptom has no matching procedure in the supplied context.

  • Compare the single source ID with the two symptoms in the printed question.
  • Identify the symptom that has no runbook-backed procedure in the answer.

That is the failure you wanted to expose. Your baseline now proves that a fluent answer can still have incomplete evidence.

Does the script stop before printing?

  • Confirm .venv remains active in the same terminal.
  • Confirm GEMINI_API_KEY remains available to the current shell without printing its value.
  • Check the active request limits in Google AI Studio if the provider declines the model request.

Help me debug the baseline

Your finished baseline.py should match the complete reference below.

✔️ Awesome, I've got everything!

Your baseline retrieves one source for the compound question. The terminal output now exposes the resulting evidence gap.

ⓧ I'd like to double check the full code

from langchain_google_genai import ChatGoogleGenerativeAI

from retrieval import search_runbooks, source_ids

MODEL = ChatGoogleGenerativeAI(
    model="gemini-3.8-flash",
    temperature=1.0,
    max_retries=2,
)

COMPOUND_QUESTION = (
    "The API is returning 503 errors and the database connection pool is exhausted. "
    "What should I do?"
)


def message_text(response: object) -> str:
    text = getattr(response, "text", None)
    return text if isinstance(text, str) else str(getattr(response, "content", response))


def baseline_answer(question: str) -> dict[str, object]:
    documents = search_runbooks(question, k=1)
    context = "\n\n".join(documents)
    response = MODEL.invoke(
        [
            (
                "system",
                "Answer only from the supplied runbook context. State when the context is incomplete.",
            ),
            ("human", f"Context:\n{context}\n\nQuestion:\n{question}"),
        ]
    )
    return {
        "question": question,
        "answer": message_text(response),
        "sources": source_ids(documents),
    }


if __name__ == "__main__":
    result = baseline_answer(COMPOUND_QUESTION)
    print(f"Question: {result['question']}")
    print(f"Sources: {result['sources']}")
    print(f"Answer:\n{result['answer']}")

This reference combines the cloud model configuration with the fixed top-one retrieval flow plus the executable compound-incident test.

Your top-one baseline now exposes the limitation in terminal output. Next, you will build a workflow that reviews its evidence coverage before accepting the answer.

Orchestrate a Corrective Agentic RAG API

Your top-one RAG baseline exposed the exact weakness you needed to see. The script cannot inspect incomplete evidence or recover from it.

You will now use LangGraph to give planning, retrieval, response, and review their own responsibilities. You will expose the resulting workflow through FastAPI so every request returns an inspectable route.

In this step, get ready to:
  • Define typed shared state for the agent nodes.
  • Build a conditional graph with one bounded retrieval retry.
  • Expose the compiled workflow through validated API endpoints.
Define the shared state and agent nodes

A graph needs one shared record that every node can read or update. The additive trace preserves each decision so you can inspect the complete route after a request finishes.

  • Return to Visual Studio Code from earlier.
  • Create `app.py` beside `baseline.py` using the new-file icon in the file sidebar.
  • Add the imports, model configuration, and shared state by pasting this code into `app.py`:
from operator import add
from time import perf_counter
from typing import Annotated, Literal, TypedDict

from fastapi import FastAPI, HTTPException
from langchain_google_genai import ChatGoogleGenerativeAI
from langgraph.graph import END, START, StateGraph
from pydantic import BaseModel, Field

from retrieval import search_runbooks, source_ids

MODEL = ChatGoogleGenerativeAI(
    model="gemini-3.8-flash",
    temperature=1.0,
    max_retries=2,
)


class FlowState(TypedDict, total=False):
    question: str
    route: Literal["direct", "retrieve"]
    documents: list[str]
    sources: list[str]
    answer: str
    review: str
    needs_retry: bool
    attempts: int
    trace: Annotated[list[str], add]

What does this foundation do?

  • `MODEL` configures the same cloud model settings as your baseline.
  • `FlowState` defines the values that can move through the graph.
  • `Annotated[list[str], add]` appends new trace entries instead of replacing earlier entries.
  • `attempts` lets the workflow enforce its two-attempt limit.
  • Save `app.py`.

Your new file now holds the contract that every agent node follows.

Seeing import warnings?

Confirm that Visual Studio Code is using the active `.venv` from earlier. Also check that `app.py` sits beside `retrieval.py`.

Help me resolve the imports in app.py.

The API models define which input is accepted and which output fields are guaranteed. Pydantic applies those contracts before the response leaves the service.

  • Add the request model, response model, and text helper below `FlowState` by pasting this code:
class AskRequest(BaseModel):
    question: str = Field(min_length=3, max_length=500)


class AskResponse(BaseModel):
    answer: str
    sources: list[str]
    trace: list[str]
    attempts: int
    latency_ms: float


def message_text(response: object) -> str:
    text = getattr(response, "text", None)
    return text if isinstance(text, str) else str(getattr(response, "content", response))

How do these models protect the API?

  • `AskRequest` limits questions to between 3 and 500 characters.
  • `AskResponse` declares every field returned to an API caller.
  • `message_text` extracts text consistently from the model response.
  • Save `app.py`.

The file now contains an input contract and an output contract for the service.

Seeing a type or indentation warning?

Keep both API classes aligned at the left edge. Keep their field declarations indented by four spaces.

Help me check these Pydantic models.

The planner decides whether a question needs runbook evidence. Exact greetings use a deterministic shortcut while other questions receive a model-backed routing decision.

  • Add the planner below `message_text` by pasting this function:
def plan_question(state: FlowState) -> dict[str, object]:
    normalized = state["question"].strip().lower().rstrip("!?.")
    if normalized in {"hello", "hi", "hey"}:
        return {"route": "direct", "trace": ["plan:direct"]}

    response = MODEL.invoke(
        [
            (
                "system",
                "You are a routing agent. Return exactly RETRIEVE for operational, incident, "
                "runbook, outage, database, disk, or API questions. Return exactly DIRECT for "
                "greetings and non-operational conversation.",
            ),
            ("human", state["question"]),
        ]
    )
    decision = message_text(response).upper()
    route: Literal["direct", "retrieve"] = (
        "direct" if "DIRECT" in decision and "RETRIEVE" not in decision else "retrieve"
    )
    return {"route": route, "trace": [f"plan:{route}"]}

How does the planner choose a route?

  • The normalized greeting check avoids spending a routing request on three common greetings.
  • The routing prompt restricts the model to two possible decisions.
  • The fallback chooses retrieval unless the response contains an unambiguous direct decision.
  • The selected route becomes the first trace entry.
  • Save `app.py`.

Your workflow now has a planner that records either `plan:direct` or `plan:retrieve`.

Planner function looks misaligned?

Check that the prompt fragments remain inside the same tuple. Also confirm that the final `return` stays inside `plan_question`.

Help me inspect the planner function.

A direct route still needs a useful response. This node handles greetings without loading runbook evidence.

  • Add the direct-response node below `plan_question` by pasting this function:
def direct_answer(state: FlowState) -> dict[str, object]:
    response = MODEL.invoke(
        [
            ("system", "Reply briefly and direct operational questions to the runbook workflow."),
            ("human", state["question"]),
        ]
    )
    return {
        "answer": message_text(response),
        "sources": [],
        "attempts": 0,
        "trace": ["direct_answer"],
    }

What does the direct node return?

  • The model produces a brief conversational response.
  • The empty source list shows that no runbook evidence was retrieved.
  • The zero attempt count confirms that retrieval never ran.
  • Save `app.py`.

The direct path now has a response format that matches the shared state.

Direct node missing a field?

Compare the four returned keys with `answer`, `sources`, `attempts`, and `trace` in the snippet.

Help me check the direct node.

Corrective retrieval needs different breadth on each attempt. The first attempt preserves the baseline's top-one behavior while the second attempt broadens the evidence set.

  • Add the retrieval node below `direct_answer` by pasting this function:
def retrieve(state: FlowState) -> dict[str, object]:
    attempt = state.get("attempts", 0) + 1
    k = 1 if attempt == 1 else 3
    documents = search_runbooks(state["question"], k=k)
    return {
        "documents": documents,
        "sources": source_ids(documents),
        "attempts": attempt,
        "trace": [f"retrieve:k={k}"],
    }

How does retrieval broaden?

  • The node increments `attempts` each time it runs.
  • Attempt one retrieves one document with `k=1`.
  • Attempt two retrieves three documents with `k=3`.
  • The trace records the exact retrieval breadth.
  • Save `app.py`.

The retrieval node can now repeat once with a broader search.

Retrieval names not resolving?

Confirm that `search_runbooks` and `source_ids` are imported from the unchanged `retrieval.py` file.

Help me reconnect the retrieval node.

The response node turns retrieved documents into an action plan. It also appends citations from the retrieved source IDs so citation presence does not depend on model formatting.

  • Add the response node below `retrieve` by pasting this function:
def respond(state: FlowState) -> dict[str, object]:
    context = "\n\n".join(state["documents"])
    response = MODEL.invoke(
        [
            (
                "system",
                "You are an incident-response agent. Answer only from the supplied runbooks. "
                "Give a short action plan, do not invent missing steps, and cite source IDs in "
                "square brackets such as [api-503].",
            ),
            ("human", f"Runbooks:\n{context}\n\nQuestion:\n{state['question']}"),
        ]
    )
    citations = " ".join(f"[{source}]" for source in state["sources"])
    answer = f"{message_text(response)}\n\nSources consulted: {citations}"
    return {"answer": answer, "trace": ["respond"]}

How is the answer grounded?

  • The node joins only the documents returned by semantic retrieval.
  • The system instruction limits the answer to those runbooks.
  • The final answer receives a deterministic citation for every retrieved source.
  • The trace records that response generation ran.
  • Save `app.py`.

Generated answers now carry an explicit list of consulted runbooks.

Citation line missing?

Confirm that the `citations` line reads from `state["sources"]`. Also check that `answer` appends `Sources consulted:` after the model text.

Help me restore deterministic citations.

The reviewer decides whether the current answer covers every symptom. A deterministic compound-question guard ensures that one retrieved source cannot pass for a question containing multiple incident parts.

  • Add the reviewer below `respond` by pasting this function:
def review_answer(state: FlowState) -> dict[str, object]:
    response = MODEL.invoke(
        [
            (
                "system",
                "You are a grounding reviewer. Return exactly PASS when the evidence and answer "
                "cover every symptom in the question. Return exactly RETRY when evidence is "
                "missing for any symptom.",
            ),
            (
                "human",
                f"Question:\n{state['question']}\n\nEvidence:\n"
                f"{'\n\n'.join(state['documents'])}\n\nAnswer:\n{state['answer']}",
            ),
        ]
    )
    verdict = message_text(response).upper()
    compound_question = any(
        marker in state["question"].lower() for marker in (" and ", " also ", " plus ")
    )
    deterministic_gap = compound_question and len(state["sources"]) < 2
    needs_retry = state["attempts"] < 2 and (
        "RETRY" in verdict or deterministic_gap
    )
    review = "retry" if needs_retry else "pass"
    return {
        "review": review,
        "needs_retry": needs_retry,
        "trace": [f"review:{review}"],
    }

Why combine two review signals?

  • The model reviews whether the evidence covers every symptom.
  • The deterministic guard detects compound wording with fewer than two sources.
  • The attempt check prevents a third retrieval.
  • The final decision records either `review:retry` or `review:pass`.
  • Save `app.py`.

All five agent responsibilities now exist as separate functions with shared state updates.

Reviewer block showing a syntax error?

Check the parentheses around `needs_retry`. Also confirm that each returned field remains inside the final dictionary.

Help me inspect the reviewer node.

Wire the corrective graph

The node functions define responsibilities while edges define control flow. Conditional edges send greetings directly to an answer or send operational questions through retrieval and review.

  • Add the two route functions below `review_answer` by pasting this code:
def route_after_plan(state: FlowState) -> Literal["direct", "retrieve"]:
    return state["route"]


def route_after_review(state: FlowState) -> Literal["retry", "done"]:
    return "retry" if state["needs_retry"] else "done"

What do the route functions return?

  • `route_after_plan` returns the decision stored by the planner.
  • `route_after_review` converts the retry flag into a graph branch name.
  • Save `app.py`.

The graph now has two small decision functions for its conditional edges.

Route return type not matching?

Check the quoted route values carefully. They must match the branch keys used by the graph.

Help me check the route functions.

The graph starts with planning and ends after either a direct response or a passing review. A retry loops back to retrieval instead of restarting the whole workflow.

  • Add the graph builder below the route functions by pasting this function:
def build_workflow():
    builder = StateGraph(FlowState)
    builder.add_node("plan", plan_question)
    builder.add_node("direct", direct_answer)
    builder.add_node("retrieve", retrieve)
    builder.add_node("respond", respond)
    builder.add_node("review", review_answer)

    builder.add_edge(START, "plan")
    builder.add_conditional_edges(
        "plan",
        route_after_plan,
        {"direct": "direct", "retrieve": "retrieve"},
    )
    builder.add_edge("direct", END)
    builder.add_edge("retrieve", "respond")
    builder.add_edge("respond", "review")
    builder.add_conditional_edges(
        "review",
        route_after_review,
        {"retry": "retrieve", "done": END},
    )
    return builder.compile()

How does the graph recover?

  • Every request begins at the `plan` node.
  • The direct branch ends after `direct_answer`.
  • The retrieval branch moves through `retrieve`, `respond`, and `review`.
  • A retry decision returns to `retrieve` for the broader second attempt.
  • Save `app.py`.

The control flow now contains a bounded correction loop instead of a single retrieve-then-generate path.

Graph branch not connecting?

Confirm that every node name in an edge matches a name passed to `add_node`. Check the two conditional mapping dictionaries for the same route values.

Help me trace the graph connections.

Compiling turns the graph definition into an executable workflow. The wrapper measures elapsed cloud response time around one complete graph invocation.

  • Add the compiled workflow and execution wrapper below `build_workflow` by pasting this code:
WORKFLOW = build_workflow()


def run_copilot(question: str) -> dict[str, object]:
    started = perf_counter()
    state = WORKFLOW.invoke(
        {"question": question, "attempts": 0, "trace": []},
        config={"recursion_limit": 12},
    )
    return {
        "answer": state["answer"],
        "sources": state.get("sources", []),
        "trace": state["trace"],
        "attempts": state.get("attempts", 0),
        "latency_ms": round((perf_counter() - started) * 1000, 2),
    }

What does the wrapper add?

  • `WORKFLOW` stores the compiled graph.
  • `run_copilot` supplies the initial attempt count and empty trace.
  • The recursion limit protects the invocation from an unintended loop.
  • `latency_ms` records the observed duration of the full workflow.
  • Save `app.py`.

Your corrective graph can now be invoked through one Python function.

Workflow compilation failing?

Confirm that `WORKFLOW = build_workflow()` appears after every node and route function. A missing function above that line prevents compilation.

Help me diagnose workflow compilation.

Expose and verify the API

The final layer turns `run_copilot` into a local web service. Request validation protects the workflow while the response model keeps every operational result inspectable.

  • Add the application and endpoint definitions below `run_copilot` by pasting this code:
app = FastAPI(
    title="Cloud Agentic RAG Incident-Response Copilot",
    version="1.0.0",
)


@app.get("/health")
def health() -> dict[str, str]:
    return {"status": "ok"}


@app.post("/ask", response_model=AskResponse)
def ask(payload: AskRequest) -> AskResponse:
    question = payload.question.strip()
    if not question:
        raise HTTPException(status_code=422, detail="Question cannot be blank")
    return AskResponse(**run_copilot(question))

What do the endpoints expose?

  • `GET /health` confirms that the service is responding.
  • `POST /ask` validates the incoming question with `AskRequest`.
  • Blank questions receive an HTTP validation error.
  • Successful requests return the fields declared by `AskResponse`.
  • Save `app.py`.

Your graph is now wrapped in a validated local API.

Endpoint decorators showing errors?

Confirm that `app = FastAPI(...)` appears before both endpoint decorators. Also check that each route starts with a forward slash.

Help me check the API layer.

Use these tabs to confirm that your complete file matches the cumulative implementation.

✔️ Awesome, I've got everything!

Great. Double-check that `app.py` is saved before starting the service.

ⓧ I'd like to double check the full code

from operator import add
from time import perf_counter
from typing import Annotated, Literal, TypedDict

from fastapi import FastAPI, HTTPException
from langchain_google_genai import ChatGoogleGenerativeAI
from langgraph.graph import END, START, StateGraph
from pydantic import BaseModel, Field

from retrieval import search_runbooks, source_ids

MODEL = ChatGoogleGenerativeAI(
    model="gemini-3.8-flash",
    temperature=1.0,
    max_retries=2,
)


class FlowState(TypedDict, total=False):
    question: str
    route: Literal["direct", "retrieve"]
    documents: list[str]
    sources: list[str]
    answer: str
    review: str
    needs_retry: bool
    attempts: int
    trace: Annotated[list[str], add]


class AskRequest(BaseModel):
    question: str = Field(min_length=3, max_length=500)


class AskResponse(BaseModel):
    answer: str
    sources: list[str]
    trace: list[str]
    attempts: int
    latency_ms: float


def message_text(response: object) -> str:
    text = getattr(response, "text", None)
    return text if isinstance(text, str) else str(getattr(response, "content", response))


def plan_question(state: FlowState) -> dict[str, object]:
    normalized = state["question"].strip().lower().rstrip("!?.")
    if normalized in {"hello", "hi", "hey"}:
        return {"route": "direct", "trace": ["plan:direct"]}

    response = MODEL.invoke(
        [
            (
                "system",
                "You are a routing agent. Return exactly RETRIEVE for operational, incident, "
                "runbook, outage, database, disk, or API questions. Return exactly DIRECT for "
                "greetings and non-operational conversation.",
            ),
            ("human", state["question"]),
        ]
    )
    decision = message_text(response).upper()
    route: Literal["direct", "retrieve"] = (
        "direct" if "DIRECT" in decision and "RETRIEVE" not in decision else "retrieve"
    )
    return {"route": route, "trace": [f"plan:{route}"]}


def direct_answer(state: FlowState) -> dict[str, object]:
    response = MODEL.invoke(
        [
            ("system", "Reply briefly and direct operational questions to the runbook workflow."),
            ("human", state["question"]),
        ]
    )
    return {
        "answer": message_text(response),
        "sources": [],
        "attempts": 0,
        "trace": ["direct_answer"],
    }


def retrieve(state: FlowState) -> dict[str, object]:
    attempt = state.get("attempts", 0) + 1
    k = 1 if attempt == 1 else 3
    documents = search_runbooks(state["question"], k=k)
    return {
        "documents": documents,
        "sources": source_ids(documents),
        "attempts": attempt,
        "trace": [f"retrieve:k={k}"],
    }


def respond(state: FlowState) -> dict[str, object]:
    context = "\n\n".join(state["documents"])
    response = MODEL.invoke(
        [
            (
                "system",
                "You are an incident-response agent. Answer only from the supplied runbooks. "
                "Give a short action plan, do not invent missing steps, and cite source IDs in "
                "square brackets such as [api-503].",
            ),
            ("human", f"Runbooks:\n{context}\n\nQuestion:\n{state['question']}"),
        ]
    )
    citations = " ".join(f"[{source}]" for source in state["sources"])
    answer = f"{message_text(response)}\n\nSources consulted: {citations}"
    return {"answer": answer, "trace": ["respond"]}


def review_answer(state: FlowState) -> dict[str, object]:
    response = MODEL.invoke(
        [
            (
                "system",
                "You are a grounding reviewer. Return exactly PASS when the evidence and answer "
                "cover every symptom in the question. Return exactly RETRY when evidence is "
                "missing for any symptom.",
            ),
            (
                "human",
                f"Question:\n{state['question']}\n\nEvidence:\n"
                f"{'\n\n'.join(state['documents'])}\n\nAnswer:\n{state['answer']}",
            ),
        ]
    )
    verdict = message_text(response).upper()
    compound_question = any(
        marker in state["question"].lower() for marker in (" and ", " also ", " plus ")
    )
    deterministic_gap = compound_question and len(state["sources"]) < 2
    needs_retry = state["attempts"] < 2 and (
        "RETRY" in verdict or deterministic_gap
    )
    review = "retry" if needs_retry else "pass"
    return {
        "review": review,
        "needs_retry": needs_retry,
        "trace": [f"review:{review}"],
    }


def route_after_plan(state: FlowState) -> Literal["direct", "retrieve"]:
    return state["route"]


def route_after_review(state: FlowState) -> Literal["retry", "done"]:
    return "retry" if state["needs_retry"] else "done"


def build_workflow():
    builder = StateGraph(FlowState)
    builder.add_node("plan", plan_question)
    builder.add_node("direct", direct_answer)
    builder.add_node("retrieve", retrieve)
    builder.add_node("respond", respond)
    builder.add_node("review", review_answer)

    builder.add_edge(START, "plan")
    builder.add_conditional_edges(
        "plan",
        route_after_plan,
        {"direct": "direct", "retrieve": "retrieve"},
    )
    builder.add_edge("direct", END)
    builder.add_edge("retrieve", "respond")
    builder.add_edge("respond", "review")
    builder.add_conditional_edges(
        "review",
        route_after_review,
        {"retry": "retrieve", "done": END},
    )
    return builder.compile()


WORKFLOW = build_workflow()


def run_copilot(question: str) -> dict[str, object]:
    started = perf_counter()
    state = WORKFLOW.invoke(
        {"question": question, "attempts": 0, "trace": []},
        config={"recursion_limit": 12},
    )
    return {
        "answer": state["answer"],
        "sources": state.get("sources", []),
        "trace": state["trace"],
        "attempts": state.get("attempts", 0),
        "latency_ms": round((perf_counter() - started) * 1000, 2),
    }


app = FastAPI(
    title="Cloud Agentic RAG Incident-Response Copilot",
    version="1.0.0",
)


@app.get("/health")
def health() -> dict[str, str]:
    return {"status": "ok"}


@app.post("/ask", response_model=AskResponse)
def ask(payload: AskRequest) -> AskResponse:
    question = payload.question.strip()
    if not question:
        raise HTTPException(status_code=422, detail="Question cannot be blank")
    return AskResponse(**run_copilot(question))

What should the full file contain?

This reference combines the state, nodes, graph, execution wrapper, and API layer in their final order. Compare it with your saved `app.py` before continuing.

Uvicorn loads the `app` object and serves it locally. Leave this terminal running while you test both endpoints.

Before you start the server, do you expect the application to import cleanly or stop before startup?

  • Start the local API from the active virtual environment by running this command:
uvicorn app:app --reload

What does this command do?

The application target uses the `<module>:<attribute>` pattern. Reload mode restarts the development server when `app.py` changes.

A clean startup confirms that Python imported the graph and compiled `WORKFLOW`. The server remains active while it waits for local requests.

Server not starting?

Confirm that the terminal still shows the active `.venv`. Check the first reported file and line in any traceback before changing code.

Help me debug the Uvicorn startup error.

  • Open `http://127.0.0.1:8000/docs` in your browser.
  • Expand `GET /health`.
  • Send a request from the expanded health endpoint.

You should receive `{"status":"ok"}`. That confirms the local API is accepting requests.

That is the service boundary working. Your local application can now serve the compiled graph without downloading a model.

Before you send the compound question, do you think the trace will stop after its first retrieval or loop through retrieval again?

  • Expand `POST /ask`.
  • Enable request editing in the expanded endpoint.
  • Enter `The API is returning 503 errors and the database connection pool is exhausted. What should I do?` as the question.
  • Send the request.

The trace should contain `retrieve:k=1`, `review:retry`, and `retrieve:k=3` in that order. The response should also show two or more source identifiers.

You should see both `api-503` and `db-pool` among the returned sources. The final answer should cite the evidence used for both symptoms.

What did the retry prove?

The first retrieval recreated the baseline's one-document limitation. The reviewer detected the evidence gap and routed the same shared state through broader retrieval.

The second review ended the graph after the expanded evidence covered the compound incident. The trace makes that correction visible instead of hiding it inside one model call.

Missing the corrective retry?

Confirm that the question still contains the word `and`. Check that `review_answer` calculates `deterministic_gap` from fewer than two sources.

Also confirm that the retry branch maps `retry` back to `retrieve`. Help me debug the missing retry.

Your API now exposes a planner, evidence retrieval, grounded generation, and corrective review through one endpoint. Next up, you will measure those behaviors across a repeatable evaluation dataset.

Measure Retrieval and Agent Behavior

Your Agentic RAG workflow now recovers from incomplete evidence with a bounded retry. A polished demo still cannot prove that this behavior stays reliable.

This step turns three representative questions into a golden evaluation dataset. The resulting report records observed evidence from live cloud runs.

In this step, get ready to:
  • Define golden cases for compound retrieval, single-runbook retrieval, and direct routing.
  • Calculate retrieval, citation, route, and latency results from live workflow runs.
  • Document the architecture, demo workflow, privacy constraints, and design decisions.
Define the golden cases

A golden case pairs a known question with the evidence or route that should appear. The three cases cover a compound incident, a disk incident, and a greeting.

  • Switch back to Visual Studio Code from the previous step.
  • Use the file sidebar to create evaluate.py inside the same project folder as app.py.
  • Add the imports and three golden cases by pasting this code into evaluate.py:
import json

from app import run_copilot

CASES = [
    {
        "name": "compound-api-and-database",
        "question": (
            "The API is returning 503 errors and the database connection pool is exhausted. "
            "What should I do?"
        ),
        "expected_sources": {"api-503", "db-pool"},
        "expected_route": "plan:retrieve",
    },
    {
        "name": "disk-capacity",
        "question": "What should I check when disk usage is critically high?",
        "expected_sources": {"disk-capacity"},
        "expected_route": "plan:retrieve",
    },
    {
        "name": "greeting",
        "question": "Hello",
        "expected_sources": set(),
        "expected_route": "plan:direct",
    },
]

What do these cases represent?

  • The compound case expects both api-503 and db-pool from the corrective retrieval route.
  • The disk case expects disk-capacity from a standard retrieval route.
  • The greeting case expects an empty source set from the direct route.
  • Save evaluate.py.
  • Confirm the editor shows the three case names ending with greeting.
  • Add the citation helper below CASES by pasting this code:
def has_citation(answer: str, source: str) -> bool:
    return f"[{source}]" in answer or f"[source:{source}]" in answer

How are citations checked?

The helper accepts the source-label formats already produced by the project. This turns citation presence into a repeatable boolean check.

  • Save evaluate.py.
  • Confirm has_citation() appears directly below the case list.

Do the expected cases look different?

Check each case name against the code block. Make sure the expected source values match the source identifiers in data/runbooks.json.

Check that greeting uses set() because the direct route consults no runbook.

Help me compare my golden cases with the expected evaluation cases.

Calculate quality and latency metrics

The evaluator calls the existing run_copilot() function for every case. Each response becomes evidence about retrieval, grounding, routing, and observed cloud latency.

  • Add the evaluation loop below has_citation() by pasting this code:
def evaluate() -> dict[str, object]:
    results = []
    for case in CASES:
        response = run_copilot(case["question"])
        expected_sources = case["expected_sources"]
        retrieval_ok = expected_sources.issubset(set(response["sources"]))
        citation_ok = all(
            has_citation(response["answer"], source) for source in expected_sources
        )
        route_ok = case["expected_route"] in response["trace"]
        results.append(
            {
                "name": case["name"],
                "retrieval_ok": retrieval_ok,
                "citation_ok": citation_ok,
                "route_ok": route_ok,
                "latency_ms": response["latency_ms"],
                "sources": response["sources"],
                "trace": response["trace"],
            }
        )

What does the evaluation loop measure?

  • retrieval_ok confirms that every expected source appears in the workflow response.
  • citation_ok confirms that the answer cites every expected source.
  • route_ok confirms that the trace contains the expected planner decision.
  • latency_ms preserves the measured cloud response time from run_copilot().
  • Save evaluate.py.
  • Confirm each result records its name, checks, latency, sources, and trace.
  • Add the summary calculation and executable entry point below the loop by pasting this code:
    total = len(results)
    summary = {
        "retrieval_pass_rate": sum(item["retrieval_ok"] for item in results) / total,
        "citation_pass_rate": sum(item["citation_ok"] for item in results) / total,
        "route_pass_rate": sum(item["route_ok"] for item in results) / total,
        "average_latency_ms": round(
            sum(item["latency_ms"] for item in results) / total,
            2,
        ),
    }
    return {"cases": results, "summary": summary}


if __name__ == "__main__":
    print(json.dumps(evaluate(), indent=2))

How is the report assembled?

The summary converts the boolean checks into observed pass rates. It also averages the measured latency from the three live runs.

json.dumps() prints the report as readable JSON. This makes the evidence easy to save or process with another tool.

  • Save evaluate.py.

✔️ Awesome, I've got everything!

Your evaluation harness is complete. Keep evaluate.py saved before you run it.

ⓧ I'd like to double check the full code

import json

from app import run_copilot

CASES = [
    {
        "name": "compound-api-and-database",
        "question": (
            "The API is returning 503 errors and the database connection pool is exhausted. "
            "What should I do?"
        ),
        "expected_sources": {"api-503", "db-pool"},
        "expected_route": "plan:retrieve",
    },
    {
        "name": "disk-capacity",
        "question": "What should I check when disk usage is critically high?",
        "expected_sources": {"disk-capacity"},
        "expected_route": "plan:retrieve",
    },
    {
        "name": "greeting",
        "question": "Hello",
        "expected_sources": set(),
        "expected_route": "plan:direct",
    },
]


def has_citation(answer: str, source: str) -> bool:
    return f"[{source}]" in answer or f"[source:{source}]" in answer


def evaluate() -> dict[str, object]:
    results = []
    for case in CASES:
        response = run_copilot(case["question"])
        expected_sources = case["expected_sources"]
        retrieval_ok = expected_sources.issubset(set(response["sources"]))
        citation_ok = all(
            has_citation(response["answer"], source) for source in expected_sources
        )
        route_ok = case["expected_route"] in response["trace"]
        results.append(
            {
                "name": case["name"],
                "retrieval_ok": retrieval_ok,
                "citation_ok": citation_ok,
                "route_ok": route_ok,
                "latency_ms": response["latency_ms"],
                "sources": response["sources"],
                "trace": response["trace"],
            }
        )

    total = len(results)
    summary = {
        "retrieval_pass_rate": sum(item["retrieval_ok"] for item in results) / total,
        "citation_pass_rate": sum(item["citation_ok"] for item in results) / total,
        "route_pass_rate": sum(item["route_ok"] for item in results) / total,
        "average_latency_ms": round(
            sum(item["latency_ms"] for item in results) / total,
            2,
        ),
    }
    return {"cases": results, "summary": summary}


if __name__ == "__main__":
    print(json.dumps(evaluate(), indent=2))

How to use this reference

Compare this reference with your saved file from top to bottom. Check the indentation inside evaluate() carefully because it controls which values belong to each case.

This suite makes several calls through the existing cloud workflow. A short pause is normal while the model and embedding requests complete.

Before you run the suite, which case do you expect to bypass retrieval?

  • Run the completed evaluation harness from the project terminal with this command:
python evaluate.py

What should the first report contain?

You should see a cases array containing three result objects. Each object includes its observed latency, source identifiers, and execution trace.

You should also see a summary object containing the four aggregate metrics. The values come from this run of the cloud workflow.

The greeting case should show plan:direct in its trace. The operational cases should show plan:retrieve.

Evaluation run interrupted?

If the import fails, confirm evaluate.py sits beside app.py in the same folder.

If a cloud request fails, confirm the virtual environment from earlier is active. Confirm the current terminal session still has access to the Gemini API key.

Help me debug my evaluation harness.

Document the evidence and design

A README turns the working service into a reproducible technical story. It connects the measured behavior to the architecture decisions behind it.

  • Use the file sidebar in Visual Studio Code to create README.md inside the project folder.
  • Add the project overview, architecture, and stack by pasting this content into README.md:
# Cloud Agentic RAG Incident-Response Copilot

A lightweight Applied AI service that keeps its API, graph, synthetic runbooks, and vector store on your Mac while using the Gemini API for embeddings and generated responses.

## Architecture

```mermaid
flowchart LR
    A[FastAPI request] --> B{Planner}
    B -->|Direct| C[Direct response]
    B -->|Retrieve| D[Gemini embedding plus local vector search]
    D --> E[Response agent]
    E --> F{Reviewer}
    F -->|Pass| G[API response]
    F -->|Retry once| D
```

## Stack

- Python and FastAPI
- LangGraph orchestration
- Gemini 3.8 Flash generation
- Gemini Embedding 2 vectors
- LangChain InMemoryVectorStore
- Pydantic request and response validation

What does the architecture show?

The diagram follows a FastAPI request through planning, retrieval, response generation, and review. The retry edge makes the corrective loop visible.

The stack list separates local application components from cloud inference. This makes the system boundary clear during a demo.

  • Save README.md.
  • Confirm the file shows the Architecture and Stack headings.
  • Add the reproducible commands and recovery explanation below the stack by pasting this content:
## Run

Create a Gemini API key in Google AI Studio, then run:

```bash
python -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt
export GEMINI_API_KEY="your-api-key-here"
uvicorn app:app --reload
```

Open `http://127.0.0.1:8000/docs` and execute `POST /ask`.

## Evaluation

```bash
python evaluate.py
```

The evaluation reports expected-source retrieval, citation presence, route behavior, and observed cloud latency. Run it yourself and publish the actual JSON output. Do not claim results you have not measured.

## Designed failure and recovery

The baseline retrieves one runbook for a question containing two incident symptoms. The Agentic RAG workflow reviews that incomplete evidence, broadens retrieval from one to three documents, and caps the process at two attempts.

Why include reproducible commands?

The run section gives another developer a direct path from setup to the interactive API. The evaluation section gives them a separate command for reproducing your evidence.

The recovery section captures the designed baseline failure. It also explains how the graph broadens retrieval within a fixed limit.

  • Save README.md.
  • Confirm the file shows the Run and Evaluation headings.
  • Add the privacy limitations and interview prompts at the end by pasting this content:
## Privacy and limitations

- The vector store is in-memory and rebuilt at process startup.
- The knowledge base is small and synthetic.
- Questions and runbook text are sent to the Gemini API. Free Tier content may be used to improve Google's products, so do not substitute confidential incident data.
- Model behavior and network latency vary between runs.
- The heuristic for compound questions is intentionally simple.
- This is a portfolio service, not an autonomous production incident responder.

## Interview talking points

- Why semantic retrieval is different from keyword matching
- Why the baseline fails on a compound question
- How LangGraph state and conditional edges implement bounded correction
- Why citations are appended as a deterministic output guardrail
- How the cloud inference boundary reduces local hardware requirements but changes privacy and availability trade-offs

Why document the limitations?

The limitations define where the demo remains intentionally lightweight. They also disclose that synthetic questions and runbooks cross the Gemini API boundary.

The interview prompts connect the implementation to engineering trade-offs. This helps you explain the system beyond a code walkthrough.

  • Save README.md.

✔️ Awesome, I've got everything!

Your README now documents how the copilot works, how to run it, and how to present measured evidence.

ⓧ I'd like to double check the full code

# Cloud Agentic RAG Incident-Response Copilot

A lightweight Applied AI service that keeps its API, graph, synthetic runbooks, and vector store on your Mac while using the Gemini API for embeddings and generated responses.

## Architecture

```mermaid
flowchart LR
    A[FastAPI request] --> B{Planner}
    B -->|Direct| C[Direct response]
    B -->|Retrieve| D[Gemini embedding plus local vector search]
    D --> E[Response agent]
    E --> F{Reviewer}
    F -->|Pass| G[API response]
    F -->|Retry once| D
```

## Stack

- Python and FastAPI
- LangGraph orchestration
- Gemini 3.8 Flash generation
- Gemini Embedding 2 vectors
- LangChain InMemoryVectorStore
- Pydantic request and response validation

## Run

Create a Gemini API key in Google AI Studio, then run:

```bash
python -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt
export GEMINI_API_KEY="your-api-key-here"
uvicorn app:app --reload
```

Open `http://127.0.0.1:8000/docs` and execute `POST /ask`.

## Evaluation

```bash
python evaluate.py
```

The evaluation reports expected-source retrieval, citation presence, route behavior, and observed cloud latency. Run it yourself and publish the actual JSON output. Do not claim results you have not measured.

## Designed failure and recovery

The baseline retrieves one runbook for a question containing two incident symptoms. The Agentic RAG workflow reviews that incomplete evidence, broadens retrieval from one to three documents, and caps the process at two attempts.

## Privacy and limitations

- The vector store is in-memory and rebuilt at process startup.
- The knowledge base is small and synthetic.
- Questions and runbook text are sent to the Gemini API. Free Tier content may be used to improve Google's products, so do not substitute confidential incident data.
- Model behavior and network latency vary between runs.
- The heuristic for compound questions is intentionally simple.
- This is a portfolio service, not an autonomous production incident responder.

## Interview talking points

- Why semantic retrieval is different from keyword matching
- Why the baseline fails on a compound question
- How LangGraph state and conditional edges implement bounded correction
- Why citations are appended as a deterministic output guardrail
- How the cloud inference boundary reduces local hardware requirements but changes privacy and availability trade-offs

How to use this reference

Compare each heading with your saved README.md. Confirm the commands, privacy statements, and interview prompts appear in the same order.

Your evaluation code and project documentation now tell the same story. The final run checks that the machine-readable evidence still comes directly from the working workflow.

Before you run the final check, what four metric names do you expect inside the summary?

  • Generate the final evaluation evidence by running this command:
python evaluate.py

What proves the evaluation works?

You should see per-case JSON followed by a summary containing retrieval_pass_rate, citation_pass_rate, route_pass_rate, and average_latency_ms.

The report preserves the sources and trace from each workflow execution. These fields let you inspect behavior behind the aggregate rates.

That completes the evidence loop. Your copilot now has reproducible checks for retrieval, citations, route behavior, and observed cloud latency.

Missing a summary metric?

Compare the final section of evaluate.py with the full-code reference. Check that the summary dictionary remains inside evaluate().

If a case reports a failed check, inspect its actual sources and trace before changing the expected values. Model-dependent variation belongs in your recorded evidence.

Help me interpret my evaluation JSON.

Secret mission

Add a New Incident Domain

Your workflow already handles three incident domains through one graph. Add a queue-backlog runbook plus a regression case to prove that the design grows through data instead of queue-specific graph code.

Clean Up Your Resources

Clean Up Your Resources

Choose whether to keep the copilot ready for future demos, pause its running session, or remove its local files and cloud credential. The guided project stays within the Gemini API Free Tier, so there are no ongoing local costs or paid cloud resources to shut down.

Resources you used:

  • The local project folder containing app.py, baseline.py, retrieval.py, evaluate.py, README.md, requirements.txt, data/runbooks.json, and .venv.
  • The Uvicorn process hosting your FastAPI service, including its ephemeral InMemoryVectorStore.
  • The GEMINI_API_KEY setting in your current Terminal session.
  • The Gemini API key managed through Google AI Studio.

Keep everything running

No cleanup action is needed. Choose this if you plan to keep testing the API or demonstrating the corrective retrieval workflow.

  • Keep the local project folder so the code, synthetic runbooks, virtual environment, and evaluation results remain available.
  • Leave the Gemini API key enabled in Google AI Studio for future model and embedding requests.
  • Continue using only synthetic runbooks because Free Tier content may be used to improve Google's products.
  • Stop Uvicorn with Control-C whenever you finish a demo.

Pause - I'll come back to this later

Shut down the local service to free its memory while keeping the project files and Gemini API key available for another session.

  • Switch back to the Terminal running Uvicorn.
  • Press Control-C to stop the service.

That pauses the local runtime. The in-process vector store is released with the Uvicorn process.

  • Remove the API key setting from the current shell by running this command:
unset GEMINI_API_KEY

What does this command do?

The command removes GEMINI_API_KEY from the current shell. The API key remains enabled in Google AI Studio for your next session.

  • Confirm the Terminal returns to its prompt without printing a value.
  • Keep the local project folder for your next demo.
  • Leave the Gemini API key enabled in Google AI Studio.

Delete - I don't want to use this again

Remove the running service, shell credential, Gemini API key, and local project files.

Check Before Deleting

Deleting credentials and files can feel final. The API key can be replaced later, but the local project folder needs a separate backup if you want to recover your work.

If Uvicorn is still running, stop its process before removing the project files.

  • Switch back to the Terminal running Uvicorn.
  • Press Control-C to stop the service.

Uvicorn stops with the in-process vector store. No local model files need removal because this project did not download a model.

  • Remove the API key setting from the current shell by running this command:
unset GEMINI_API_KEY

What does this command do?

The command removes GEMINI_API_KEY from the current shell. It does not delete the API key stored in Google AI Studio.

  • Confirm the Terminal returns to its prompt without printing a value.
  • Return to Google AI Studio from earlier.
  • Select Dashboard.
  • Select API Keys.
  • Select the Gemini API key used for this project.
  • Delete the project key from the API Keys page.
  • Complete the confirmation prompt.
  • Confirm the API Keys page no longer shows the key as active.

That's the sensitive part handled: the key can no longer authorize Gemini API requests.

  • Use Finder to locate the local project folder from earlier.
  • Delete the entire project folder through Finder.
  • Confirm the folder no longer appears in Finder.

The source files and .venv are removed together because both lived in that folder.

If You Enabled Billing Separately

Paid billing was outside this project.

  • Use Google Cloud Billing to confirm that no paid resources remain in the associated project.

Cleanup is complete. The local runtime, shell credential, cloud API key, and project files are no longer available.

Nice Work!

Nice Work!

That is the full system working! Your incident-response copilot now returns cited guidance from synthetic runbooks through a tested Agentic RAG API.

What you learned:

  • Built semantic retrieval with Gemini Embedding 2 to match incident questions with synthetic runbooks. Demonstrated the single-pass RAG limitation when top-one retrieval missed part of a compound incident.
  • Orchestrated distinct agent responsibilities through LangGraph shared state. Served the corrective workflow through a validated FastAPI endpoint. A bounded retry expands incomplete evidence before the response returns.
  • Produced a repeatable evaluation report for source retrieval. Added checks for citation presence. Verified graph routes. Recorded observed cloud latency.
  • Secret Mission: Extended the knowledge base with a queue-backlog runbook. Proved the data-driven design with a new regression case. The retrieval engine and graph topology stayed unchanged.

Ready to quiz yourself?