Build Governed Supplier Approval Agent

Build a governed Gemini agent with RAG, approvals, and policy controls.

Introduction

30 Second Summary

Supplier approvals can look straightforward until spending limits, security checks, or sanctions rules enter the picture. A confident recommendation can still expose a business to risk when it overlooks one policy.

In this project, you will build a terminal-based supplier-approval assistant using Gemini, LangGraph, and retrieval-augmented generation (RAG). It finds relevant internal policies before deterministic rules block unsafe requests or pause permitted requests for your approval.

What You'll Build

You will run a fictional supplier request in your terminal to see cited policy evidence, a guarded decision, and a human approval checkpoint before any record is written.

By the end of this project, you'll have:

  • A grounded recommendation that displays relevant policy IDs beside the model's rationale.
  • Deterministic safety gates that block requests with sanctions or security review failures before they reach the write boundary.
  • A human-controlled workflow that pauses before recording an approval. You can also inspect local records, a three-case evaluation report, and an execution trace in LangSmith.
  • Secret Mission: Expose policy retrieval through a typed read-only MCP tool. Prove its contract through an in-memory client.

Are there any prerequisites?

You need Linux with Python 3.12 and OpenAI Codex already installed. The guide walks you through creating a Gemini API key with Google AI Studio plus a payment-free LangSmith Developer account.

Before We Start

A governed agent needs a clear decision boundary before any code exists. This checkpoint defines the supplier recommendation while keeping every final approval-record write under human control.

Set Up the Free Agent Workspace

The business goal is clear. Your workspace now needs to make that goal reproducible with an isolated Python environment.

Pinned dependencies prevent version drift. An architecture check selects the compatible LanceDB package for your Linux machine.

In this step, get ready to:
  • Create a Linux workspace with pinned dependency files.
  • Store both API keys in ignored local configuration.
  • Install the dependency set that matches your machine architecture.
Build the isolated workspace

A virtual environment keeps this project's packages separate from the rest of your machine. The supplied files also pin every dependency to a known version.

  • Open your Linux terminal from the applications menu.
  • Move to your Desktop by running this command:
cd ~/Desktop

What does this command do?

The command moves your terminal to the Desktop. Your new workspace will be easy to find there.

  • Create the workspace plus both architecture-specific requirements files by running these commands:
mkdir governed-supplier-agent
cd governed-supplier-agent
cat > requirements-x86_64.txt <<'EOF'
google-genai==2.28.0
langgraph==1.2.12
lancedb-compat==0.38.0
langsmith==0.14.4
mcp==2.3.0
pydantic==2.13.5
python-dotenv==1.2.4
EOF

cat > requirements-arm64.txt <<'EOF'
google-genai==2.28.0
langgraph==1.2.12
lancedb==0.39.0
langsmith==0.14.4
mcp==2.3.0
pydantic==2.13.5
python-dotenv==1.2.4
EOF

What did this create?

  • The governed-supplier-agent folder contains every local project file.
  • The requirements-x86_64.txt file uses the LanceDB compatibility package for x86_64 machines.
  • The requirements-arm64.txt file uses the standard LanceDB package for ARM64 machines.
  • Confirm that both requirements files exist by running this command:
ls

What should I see?

You should see requirements-x86_64.txt plus requirements-arm64.txt in the output. This confirms that both installation paths are ready.

  • Create the environment template plus ignore rules by running these commands:
cat > .env.example <<'EOF'
GEMINI_API_KEY=your-api-key-here
LANGSMITH_TRACING=true
LANGSMITH_API_KEY=your-langsmith-api-key-here
LANGSMITH_PROJECT=governed-supplier-agent
EOF

cat > .gitignore <<'EOF'
.env
.venv/
__pycache__/
*.pyc
data/lancedb/
artifacts/
EOF

How do these files protect the project?

The .env.example file documents the required configuration without containing live credentials. The LANGSMITH_PROJECT value groups future traces under governed-supplier-agent.

The .gitignore file tells Git to exclude secrets plus generated local data. This keeps .env out of future commits.

  • Confirm that the hidden configuration files exist by running this command:
ls -a

What should I see now?

You should now see .env.example plus .gitignore beside both requirements files. The -a option includes hidden files in the listing.

  • Create the .venv environment plus activate it by running these commands:
python3 -m venv .venv
source .venv/bin/activate

Why isolate the environment?

The first command creates a local Python environment inside .venv. The second command makes that environment active in the current terminal.

Future installations now stay inside this workspace. That isolation protects other Python projects from dependency conflicts.

  • Check where pip resolves by running this command:
python3 -m pip --version

What confirms activation?

The printed path should contain .venv. This proves that package installation targets the isolated environment.

Virtual environment not activating?

Confirm that your terminal is inside the governed-supplier-agent folder. Check that .venv/bin/activate exists before running the activation command again.

If the path still points outside .venv, help me diagnose my virtual environment.

✔️ Awesome, I've got everything!

Your four supplied project files are ready. Keep the terminal open because the active virtual environment belongs to this shell session.

ⓧ I'd like to double check the full code

Compare each file with the complete versions below. Preserve every version pin plus every line break.

google-genai==2.28.0
langgraph==1.2.12
lancedb-compat==0.38.0
langsmith==0.14.4
mcp==2.3.0
pydantic==2.13.5
python-dotenv==1.2.4

Why this x86_64 file matters

This file pins the x86_64 dependency set. Its compatibility package still exposes the standard lancedb import used later.

google-genai==2.28.0
langgraph==1.2.12
lancedb==0.39.0
langsmith==0.14.4
mcp==2.3.0
pydantic==2.13.5
python-dotenv==1.2.4

Why this ARM64 file matters

This file pins the ARM64 dependency set. It installs the standard LanceDB distribution for that architecture.

GEMINI_API_KEY=your-api-key-here
LANGSMITH_TRACING=true
LANGSMITH_API_KEY=your-langsmith-api-key-here
LANGSMITH_PROJECT=governed-supplier-agent

Why the template stays public

This template records the required variable names with safe placeholders. Live credentials belong only in the ignored .env copy.

.env
.venv/
__pycache__/
*.pyc
data/lancedb/
artifacts/

Why these paths are ignored

These rules exclude credentials plus generated environment data. They also exclude the local database plus future approval artifacts.

Create the API credentials

The agent needs a Gemini Developer API key for model calls. It also needs a LangSmith key for future traces.

Keep usage free and data fictional

Gemini 3.8 Flash plus Gemini Embedding 2 are free of charge on the Free tier. A payment-free LangSmith Developer account stops at 5,000 base traces per month.

Use only fictional supplier data throughout this project. Gemini Free tier content is used to improve Google products.

  • Create the local credential file from the safe template by running these commands:
cp .env.example .env
ls -a

What did these commands do?

The first command creates the local .env file. The second command confirms that it now exists beside .env.example.

Creating credentials can feel risky. These keys stay local in the ignored .env file while the reusable template keeps safe placeholders.

  • Open Google AI Studio in your browser.
  • Sign in with your Google account.
  • Click Create API key.
  • Keep the project on the Free tier by leaving billing disabled.
  • Copy the generated Gemini API key.

Your Gemini credential is ready for the local environment file. Keep it in your clipboard while you create the second credential.

  • Open the LangSmith account page in a new browser tab.
  • Create a Developer account without adding a payment method.
  • Select Settings.
  • Select API Keys.
  • Choose Personal Access Tokens.
  • Click Create API Key.
  • Copy the generated LangSmith API key.
  • Return to the governed-supplier-agent workspace from earlier.
  • Open .env in your editor.
  • Replace your-api-key-here with your Gemini API key.
  • Replace your-langsmith-api-key-here with your LangSmith API key.
  • Save .env.

Your local configuration now enables tracing for governed-supplier-agent. Your live key values stay outside the supplied project files.

Unsure whether the secrets are protected?

Confirm that .gitignore contains a line with only .env. Keep the live key values out of .env.example.

If your files differ, help me review my secret-file setup without exposing any key values.

Install the matching dependency set

Your CPU architecture decides which pinned requirements file to install. Both paths provide the same project-facing lancedb import.

  • Check the machine architecture by running this command:
uname -m

How does the result choose a file?

An x86_64 result selects requirements-x86_64.txt. An aarch64 result selects requirements-arm64.txt.

The installation can take a few minutes while pip downloads the pinned packages. A quiet pause during a compiled package download is expected.

  • Select the tab that matches the architecture printed in your terminal.

x86_64

  • Install the x86_64 dependency set by running this command:
python3 -m pip install -r requirements-x86_64.txt

What does this install?

This command installs every pinned x86_64 dependency inside the active .venv environment. The dependency set includes the LanceDB compatibility distribution for the x86-64-v2 baseline.

  • Let the command finish.

You should see the terminal prompt return without an installation error. The x86_64 environment is now ready.

Dependency installation failed?

Check that uname -m printed x86_64. Confirm that your terminal prompt still shows the active .venv environment.

For an environment-specific failure, help me diagnose my x86_64 dependency installation.

aarch64

  • Install the ARM64 dependency set by running this command:
python3 -m pip install -r requirements-arm64.txt

What does this install?

This command installs every pinned ARM64 dependency inside the active .venv environment. The dependency set includes the standard LanceDB distribution.

  • Let the command finish.

You should see the terminal prompt return without an installation error. The ARM64 environment is now ready.

Dependency installation failed?

Check that uname -m printed aarch64. Confirm that your terminal prompt still shows the active .venv environment.

For an environment-specific failure, help me diagnose my ARM64 dependency installation.

Before you run the final check, which architecture should print first? Which directory should appear inside the pip path?

  • Verify the architecture plus active package location by running these commands:
uname -m
python3 -m pip --version

What should I see?

The first line should print x86_64 or aarch64. The second result should contain .venv in its path.

Together, these results prove that you selected the matching dependency file. They also prove that pip is isolated inside the project environment.

That's the workspace secured: your dependencies are pinned to the machine architecture. Your credentials are also stored outside the supplied files.

Your free agent workspace is ready. Next, you will make an intentionally ungrounded Gemini recommendation and expose the evidence gap.

Expose the Grounding Gap

Your workspace can now call Gemini from the activated virtual environment. LangSmith tracing is ready to capture the request.

However, a fluent recommendation can ignore the procurement rules that make a decision defensible. This step deliberately gives the model no policy evidence so you can see the grounding gap.

In this step, get ready to:
  • Add three fictional supplier policies plus one sample request.
  • Build a Gemini recommendation script with no policy context.
  • Run the script to expose its failed grounding check.
Add the fictional business data

The supplier request needs local business rules that we can compare against the model's answer. Keeping this data fictional also protects real supplier information from being sent to the hosted API.

  • In your project workspace's file tree, create a folder named data.

You should now see the data folder beside .env.example in your workspace.

  • Inside data, create a file named policies.json.
  • Paste the supplied policy records into policies.json:
[
  {
    "policy_id": "POL-SEC-001",
    "title": "Customer data security review",
    "text": "A supplier that handles customer data must complete a security review before approval."
  },
  {
    "policy_id": "POL-FIN-002",
    "title": "High-spend approval",
    "text": "A supplier request with annual spend of USD 50,000 or more requires human director approval before an approval record is written."
  },
  {
    "policy_id": "POL-CMP-003",
    "title": "Sanctions screening",
    "text": "A supplier may not be approved unless sanctions screening status is clear."
  }
]

What do these policies cover?

  • The POL-SEC-001 record requires a completed security review when a supplier handles customer data.
  • The POL-FIN-002 record requires director approval for annual spend of at least USD 50,000.
  • The POL-CMP-003 record blocks approval while sanctions screening remains unclear.
  • Save data/policies.json.
  • Confirm the file contains three records with distinct policy_id values.

Does the policy file look incomplete?

Check that the outer square brackets surround all three policy objects. Check that a comma separates each object.

Compare the final POL-CMP-003 record with the code block above if your editor highlights the JSON.

Help me check my policy JSON

The sample request triggers the high-spend rule while satisfying the security review requirement. Its sanctions screening is also clear.

  • Inside data, create a file named sample_request.json.
  • Paste the fictional Northstar Analytics request into sample_request.json:
{
  "request_id": "SUP-2026-001",
  "supplier_name": "Northstar Analytics",
  "annual_spend_usd": 75000,
  "handles_customer_data": true,
  "security_review_complete": true,
  "sanctions_screening": "clear"
}

What does this request test?

Northstar Analytics requests USD 75,000 in annual spend. That amount crosses the threshold in POL-FIN-002.

The request also handles customer data. Its completed security review satisfies the requirement in POL-SEC-001.

  • Save data/sample_request.json.
  • Confirm the request ID at the top of the file is SUP-2026-001.

Does the request file show a warning?

Check that the values true appear without quotation marks. Check that every property except the last one ends with a comma.

Help me check my sample request

Build the policy-blind model call

The first script gives Gemini only the supplier request. The missing policy context creates the designed obstacle for this step.

  • In your project workspace's file tree, create naive_agent.py beside .env.example.
  • Paste the runtime setup below into naive_agent.py:
import json
from pathlib import Path

from dotenv import load_dotenv

load_dotenv()

from google import genai
from langsmith import traceable

GENERATION_MODEL = "gemini-3.8-flash"

How does the runtime setup work?

  • The json module converts the request into text for the prompt.
  • The Path class gives the script a direct path to the sample request.
  • The load_dotenv() call loads your locally stored environment configuration.
  • The GENERATION_MODEL constant selects gemini-3.8-flash for the recommendation.
  • Save naive_agent.py.
  • Confirm the final visible line sets GENERATION_MODEL to gemini-3.8-flash.

Is the setup code highlighted as invalid?

Check that each import appears on its own line. Check that the model ID has matching quotation marks.

Help me fix the runtime setup

The recommendation function builds one prompt from the request. It contains no policy text or citation contract.

  • Place the recommendation function directly below the GENERATION_MODEL line by pasting this code:
@traceable(run_type="llm", name="Naive Gemini supplier recommendation")
def generate_naive_recommendation(request: dict[str, object]) -> str:
    prompt = (
        "You are a procurement assistant. Recommend approve, reject, or review "
        "for this supplier request and explain the decision briefly.\n\n"
        f"Request: {json.dumps(request)}"
    )
    with genai.Client() as client:
        interaction = client.interactions.create(
            model=GENERATION_MODEL,
            input=prompt,
        )
    return interaction.output_text or "No response returned."

How does the naive call work?

  • The @traceable decorator records the model function as an LLM run in LangSmith.
  • The prompt asks for a recommendation based only on the supplier request.
  • The genai.Client() call reads the Gemini credential from your environment.
  • The function returns the model's response text to the terminal.
  • Save naive_agent.py.
  • Confirm the function ends with the fallback text No response returned..

Does the function indentation look uneven?

Keep the prompt construction inside generate_naive_recommendation(). Keep the final return aligned with the with statement.

Help me fix the recommendation function

The entry point loads the fictional request before calling the recommendation function. It also prints a fixed check that records the missing policy evidence.

  • Place the entry point at the bottom of naive_agent.py by pasting this code:
def main() -> None:
    request = json.loads(Path("data/sample_request.json").read_text())
    print(generate_naive_recommendation(request))
    print("\nGROUNDING CHECK: FAIL - no policy IDs were supplied")


if __name__ == "__main__":
    main()

What does the entry point do?

  • The main() function reads data/sample_request.json into a Python dictionary.
  • The first print() call displays Gemini's recommendation.
  • The second print() call displays the deterministic grounding result.
  • The final condition runs main() when you execute the file directly.
  • Save naive_agent.py.
  • Confirm the final line in the file is main().

Is the entry point nested inside the function?

Align the if __name__ == "__main__": line with def main(). Indent only its following main() call.

Help me fix the entry point

✔️ Awesome, I've got everything!

Your saved script now loads the sample request and sends it to Gemini without policy context.

ⓧ I'd like to double check the full code

import json
from pathlib import Path

from dotenv import load_dotenv

load_dotenv()

from google import genai
from langsmith import traceable

GENERATION_MODEL = "gemini-3.8-flash"


@traceable(run_type="llm", name="Naive Gemini supplier recommendation")
def generate_naive_recommendation(request: dict[str, object]) -> str:
    prompt = (
        "You are a procurement assistant. Recommend approve, reject, or review "
        "for this supplier request and explain the decision briefly.\n\n"
        f"Request: {json.dumps(request)}"
    )
    with genai.Client() as client:
        interaction = client.interactions.create(
            model=GENERATION_MODEL,
            input=prompt,
        )
    return interaction.output_text or "No response returned."


def main() -> None:
    request = json.loads(Path("data/sample_request.json").read_text())
    print(generate_naive_recommendation(request))
    print("\nGROUNDING CHECK: FAIL - no policy IDs were supplied")


if __name__ == "__main__":
    main()

What should the complete script contain?

The complete file contains the runtime setup plus the recommendation function. It ends with the entry point that loads the request and prints the grounding check.

Run the grounding check

Grounding ties a model response to supplied evidence. This script has access to the supplier request but receives none of the three policy records.

  • Switch back to the terminal where .venv is active.

The request can sit quietly for a few seconds while Gemini responds. That pause means the hosted model is processing the prompt.

Before you run this, do you expect the recommendation to cite any of the three policy IDs?

  • Run the policy-blind supplier recommendation with this command:
python naive_agent.py

What should you see?

You should see a Gemini recommendation first. Its exact wording can vary between runs.

The final line should read GROUNDING CHECK: FAIL - no policy IDs were supplied.

That failure is intentional. The prompt never supplies policy IDs or policy text for the model to cite.

Did the request fail before the grounding check?

Confirm your terminal still shows the activated .venv environment. Confirm your local .env file contains a value for GEMINI_API_KEY.

Check that data/sample_request.json sits inside the data folder. Check that you ran the command from the workspace containing naive_agent.py.

Help me debug the naive agent run

You have exposed the exact weakness this project needs to solve. The model can produce a recommendation while the deterministic check proves that no policy evidence supported it.

Your policy-blind baseline is working. Next up, you will retrieve local policy evidence before asking Gemini to reason about the supplier.

Build the Local Policy Retriever

Your first Gemini call produced a fluent supplier recommendation. Its answer could not point to a single policy that supported the decision.

This step builds a local RAG retrieval layer that finds evidence before the model reasons. Gemini turns each policy into a 768-dimensional embedding. LanceDB stores those vectors locally for policy searches.

In this step, get ready to:
  • Prepare policy documents and queries for retrieval.
  • Build a local policy table from Gemini embeddings.
  • Search the table for relevant policy evidence.
Prepare policy text and embeddings

Retrieval works best when queries and documents have distinct prefixes. The prefixes tell the embedding model whether the text represents a question or a possible result.

  • Create retrieval.py inside your project folder using your code editor.
  • Add the retrieval foundation below to retrieval.py:
import argparse
import json
from pathlib import Path

from dotenv import load_dotenv

load_dotenv()

from google import genai
from google.genai import types
import lancedb
from langsmith import traceable

DB_PATH = "data/lancedb"
TABLE_NAME = "supplier_policies"
EMBEDDING_MODEL = "gemini-embedding-2"
EMBEDDING_DIMENSIONS = 768
POLICY_PATH = Path("data/policies.json")


def prepare_query(query: str) -> str:
    return f"task: search result | query: {query}"


def prepare_document(title: str, content: str) -> str:
    return f"title: {title} | text: {content}"

What does this foundation do?

  • The constants keep the database path, table name, model ID, vector size, and policy path consistent throughout the retriever.
  • The prepare_query() function applies the query prefix expected by the embedding model.
  • The prepare_document() function combines each policy title with its text using the document prefix.
  • Save retrieval.py.
  • Use your editor search to find EMBEDDING_DIMENSIONS = 768.

You should find one matching constant in the saved file.

Cannot find the embedding constant?

Check that you created retrieval.py inside the same project folder as naive_agent.py.

Compare the constant name with EMBEDDING_DIMENSIONS because Python identifiers are case-sensitive. Help me check the retrieval foundation.

The next function sends prepared text to Gemini. It requests 768-dimensional vectors so every stored policy and incoming query share the same shape.

  • Add the embedding and policy-loading functions below prepare_document() in retrieval.py:
@traceable(run_type="llm", name="Gemini policy embedding")
def embed_text(text: str) -> list[float]:
    with genai.Client() as client:
        result = client.models.embed_content(
            model=EMBEDDING_MODEL,
            contents=text,
            config=types.EmbedContentConfig(
                output_dimensionality=EMBEDDING_DIMENSIONS
            ),
        )
    return list(result.embeddings[0].values)


def _load_policies() -> list[dict[str, str]]:
    return json.loads(POLICY_PATH.read_text())

How does embedding work here?

  • The embed_text() function sends one prepared string to the embedding API.
  • The dimensionality setting produces a consistent 768-number vector for every piece of text.
  • The traceable decorator nests each embedding call inside your LangSmith trace.
  • The _load_policies() function reads the local JSON policy list.
  • Save retrieval.py.
  • Use your editor search to find def embed_text.

You should find one traced function that returns the embedding values as a list.

Seeing unresolved imports?

Confirm that your terminal still shows the activated .venv environment. Your architecture-specific requirements file already provides the package imported as lancedb.

Check that from google.genai import types appears above embed_text(). Help me diagnose the unresolved imports.

Build and search the policy table

A local vector table keeps each policy beside its embedding. Rebuilding the table refreshes every row from data/policies.json while normal searches reuse the stored table.

  • Add the index-building function below _load_policies() in retrieval.py:
@traceable(run_type="tool", name="Build supplier policy index")
def ensure_index(rebuild: bool = False):
    db = lancedb.connect(DB_PATH)
    table_names = set(db.list_tables().tables)
    if rebuild or TABLE_NAME not in table_names:
        rows = []
        for policy in _load_policies():
            rows.append(
                {
                    **policy,
                    "vector": embed_text(
                        prepare_document(policy["title"], policy["text"])
                    ),
                }
            )
        mode = "overwrite" if TABLE_NAME in table_names else "create"
        return db.create_table(TABLE_NAME, data=rows, mode=mode)
    return db.open_table(TABLE_NAME)

What does the index builder do?

  • The ensure_index() function connects to the local database under data/lancedb.
  • Each policy row keeps its original fields beside a vector generated from its title and text.
  • The rebuild option overwrites an existing supplier_policies table with fresh vectors.
  • A normal call opens the existing table so repeated searches avoid rebuilding it.
  • Save retrieval.py.
  • Use your editor search to find db.create_table(TABLE_NAME, data=rows, mode=mode).

You should find the table-creation call inside the rebuild branch.

Is the table builder incomplete?

Check that the for policy in _load_policies(): loop stays inside the rebuild branch. Confirm that the final return db.open_table(TABLE_NAME) line aligns with the branch itself.

Help me check the index-builder indentation.

The stored vectors become useful when a query is embedded with the matching query prefix. LanceDB compares that vector with the policy vectors before returning the closest records.

  • Add the policy-search function below ensure_index() in retrieval.py:
@traceable(run_type="tool", name="Search supplier policies")
def retrieve_policies(query: str, limit: int = 3) -> list[dict[str, object]]:
    table = ensure_index()
    query_vector = embed_text(prepare_query(query))
    return (
        table.search(query_vector)
        .select(["policy_id", "title", "text"])
        .limit(limit)
        .to_list()
    )

How does the search stay focused?

  • The default limit keeps each search to the three closest policy records.
  • The selection returns only policy_id, title, and text fields.
  • The final list provides ordinary Python dictionaries that later agent nodes can consume.
  • Save retrieval.py.
  • Use your editor search to find .limit(limit).

You should find the limit directly before the search results become a list.

Does the search chain look broken?

Keep the search operations inside the surrounding parentheses. Check that each chained operation begins with a full stop.

Help me repair the LanceDB search chain.

Add the query interface and verify retrieval

The retriever needs one consistent way to describe a complete supplier request. It also needs a command-line entry point that rebuilds the table and prints evidence you can inspect.

  • Add the request-query helper below retrieve_policies() in retrieval.py:
def build_request_query(request: dict[str, object]) -> str:
    return (
        f"Supplier {request['supplier_name']} requests annual spend of USD "
        f"{request['annual_spend_usd']}. Handles customer data: "
        f"{request['handles_customer_data']}. Security review complete: "
        f"{request['security_review_complete']}. Sanctions screening: "
        f"{request['sanctions_screening']}."
    )

Why build one request string?

The build_request_query() function converts every risk field into searchable prose. A consistent format helps each supplier request reach the same retrieval path.

  • Save retrieval.py.
  • Use your editor search to find sanctions_screening inside build_request_query().

You should see the sanctions status included in the generated request description.

Seeing mismatched request fields?

Compare each dictionary key with data/sample_request.json. Keep the underscores and letter casing unchanged.

Help me check the request-query fields.

  • Add the command-line entry point below build_request_query() in retrieval.py:
def main() -> None:
    parser = argparse.ArgumentParser()
    parser.add_argument("query", nargs="?", default="Which supplier policies apply?")
    parser.add_argument("--rebuild", action="store_true")
    args = parser.parse_args()

    ensure_index(rebuild=args.rebuild)
    results = retrieve_policies(args.query)
    for result in results:
        print(f"{result['policy_id']}: {result['title']}")


if __name__ == "__main__":
    main()

What does the command-line interface do?

  • The optional query argument accepts the policy question you want to search.
  • The rebuild flag refreshes the table before retrieval.
  • The output loop prints each matched policy ID beside its title.
  • Save retrieval.py.
  • Use your editor search to find if __name__ == "__main__":.

You should find the entry-point guard at the bottom of the file.

Is the entry point outside the file?

Confirm that main() remains indented beneath the entry-point guard. Keep the argument parsing inside the main() function.

Help me check the command-line entry point.

✔️ Awesome, I've got everything!

Your saved retrieval.py file now contains the complete local policy retriever.

ⓧ I'd like to double check the full code

import argparse
import json
from pathlib import Path

from dotenv import load_dotenv

load_dotenv()

from google import genai
from google.genai import types
import lancedb
from langsmith import traceable

DB_PATH = "data/lancedb"
TABLE_NAME = "supplier_policies"
EMBEDDING_MODEL = "gemini-embedding-2"
EMBEDDING_DIMENSIONS = 768
POLICY_PATH = Path("data/policies.json")


def prepare_query(query: str) -> str:
    return f"task: search result | query: {query}"


def prepare_document(title: str, content: str) -> str:
    return f"title: {title} | text: {content}"


@traceable(run_type="llm", name="Gemini policy embedding")
def embed_text(text: str) -> list[float]:
    with genai.Client() as client:
        result = client.models.embed_content(
            model=EMBEDDING_MODEL,
            contents=text,
            config=types.EmbedContentConfig(
                output_dimensionality=EMBEDDING_DIMENSIONS
            ),
        )
    return list(result.embeddings[0].values)


def _load_policies() -> list[dict[str, str]]:
    return json.loads(POLICY_PATH.read_text())


@traceable(run_type="tool", name="Build supplier policy index")
def ensure_index(rebuild: bool = False):
    db = lancedb.connect(DB_PATH)
    table_names = set(db.list_tables().tables)
    if rebuild or TABLE_NAME not in table_names:
        rows = []
        for policy in _load_policies():
            rows.append(
                {
                    **policy,
                    "vector": embed_text(
                        prepare_document(policy["title"], policy["text"])
                    ),
                }
            )
        mode = "overwrite" if TABLE_NAME in table_names else "create"
        return db.create_table(TABLE_NAME, data=rows, mode=mode)
    return db.open_table(TABLE_NAME)


@traceable(run_type="tool", name="Search supplier policies")
def retrieve_policies(query: str, limit: int = 3) -> list[dict[str, object]]:
    table = ensure_index()
    query_vector = embed_text(prepare_query(query))
    return (
        table.search(query_vector)
        .select(["policy_id", "title", "text"])
        .limit(limit)
        .to_list()
    )


def build_request_query(request: dict[str, object]) -> str:
    return (
        f"Supplier {request['supplier_name']} requests annual spend of USD "
        f"{request['annual_spend_usd']}. Handles customer data: "
        f"{request['handles_customer_data']}. Security review complete: "
        f"{request['security_review_complete']}. Sanctions screening: "
        f"{request['sanctions_screening']}."
    )


def main() -> None:
    parser = argparse.ArgumentParser()
    parser.add_argument("query", nargs="?", default="Which supplier policies apply?")
    parser.add_argument("--rebuild", action="store_true")
    args = parser.parse_args()

    ensure_index(rebuild=args.rebuild)
    results = retrieve_policies(args.query)
    for result in results:
        print(f"{result['policy_id']}: {result['title']}")


if __name__ == "__main__":
    main()

Before you run this, consider which policy ID the high-spend phrase should surface.

  • Rebuild the local index and search it by running this command:
python retrieval.py --rebuild "Which policy applies to a high-spend supplier handling customer data?"

What should you see?

The command embeds all three policies before searching the rebuilt table. You should see three policy lines that include POL-FIN-002 plus at least one additional relevant policy ID.

Those IDs prove that the query now reaches local policy evidence before any supplier decision is drafted.

Missing the high-spend policy?

Confirm that the command includes the rebuild flag. Check that data/policies.json still contains POL-FIN-002 with its high-spend policy text.

If the command fails before printing results, compare retrieval.py with the full-code tab. Help me debug the policy retriever.

That grounding gap is closed. Your supplier workflow can now retrieve policy IDs before it asks a model to explain a decision.

Your local evidence layer is ready. Next, you will place deterministic gates and human approval between that evidence and the simulated write.

Gate Requests and Require Human Approval

Your local LanceDB retriever now supplies the policy evidence that the ungrounded recommendation was missing. Retrieval makes the decision explainable.

Policy evidence cannot safely authorize a business action on its own. In this step, deterministic Python rules control the outcome while Gemini explains it. A LangGraph interrupt keeps the final write under human control.

In this step, get ready to:
  • Create a structured decision model with deterministic business gates.
  • Build a state graph that blocks unsafe requests before the write boundary.
  • Approve a permitted request through a human review interrupt.
Create the governed agent

A deterministic control plane applies business rules as ordinary Python conditions. The model produces a structured recommendation through Pydantic. It never owns the final gate status.

  • Create a new file named agent.py inside the same project folder as retrieval.py.
  • Select the second tab below to view the complete file.
  • Copy the complete code into agent.py.

✔️ I've added the file

Your governed workflow is ready for a final comparison.

  • Confirm that agent.py contains the state model, graph nodes, approval interrupt, and terminal runner.

ⓧ I'd like to double check the full code

import json
from datetime import datetime, timezone
from pathlib import Path
from typing import Literal, TypedDict

from dotenv import load_dotenv

load_dotenv()

from google import genai
from langgraph.checkpoint.memory import InMemorySaver
from langgraph.graph import END, START, StateGraph
from langgraph.types import Command, interrupt
from langsmith import traceable
from pydantic import BaseModel

from retrieval import build_request_query, retrieve_policies

GENERATION_MODEL = "gemini-3.8-flash"
ACTION_LOG = Path("artifacts/approved_suppliers.jsonl")


class DecisionDraft(BaseModel):
    recommendation: Literal["APPROVE", "REJECT", "REVIEW"]
    rationale: str
    cited_policy_ids: list[str]


class AgentState(TypedDict, total=False):
    request: dict[str, object]
    policies: list[dict[str, object]]
    gate_status: str
    draft: dict[str, object]
    human_approved: bool
    final_status: str
    action_record: str


def deterministic_gate(request: dict[str, object]) -> str:
    if str(request["sanctions_screening"]).lower() != "clear":
        return "BLOCKED"
    if bool(request["handles_customer_data"]) and not bool(
        request["security_review_complete"]
    ):
        return "BLOCKED"
    if float(request["annual_spend_usd"]) >= 50_000:
        return "DIRECTOR_REVIEW_REQUIRED"
    return "HUMAN_REVIEW_REQUIRED"


@traceable(run_type="llm", name="Gemini grounded supplier recommendation")
def generate_draft(
    request: dict[str, object],
    policies: list[dict[str, object]],
    gate_status: str,
) -> DecisionDraft:
    policy_context = "\n".join(
        f"[{policy['policy_id']}] {policy['title']}: {policy['text']}"
        for policy in policies
    )
    prompt = f"""
You are a procurement decision assistant.
Use only the supplied policy context.
The deterministic gate status is {gate_status}; do not override it.
If the gate is BLOCKED, recommend REJECT. Otherwise recommend REVIEW.
Cite only policy IDs present in the context.

Supplier request:
{json.dumps(request, indent=2)}

Policy context:
{policy_context}
""".strip()

    with genai.Client() as client:
        interaction = client.interactions.create(
            model=GENERATION_MODEL,
            input=prompt,
            response_format={
                "type": "text",
                "mime_type": "application/json",
                "schema": DecisionDraft.model_json_schema(),
            },
        )
    if not interaction.output_text:
        raise RuntimeError("Gemini returned no structured decision.")
    return DecisionDraft.model_validate_json(interaction.output_text)


def retrieve_node(state: AgentState) -> AgentState:
    query = build_request_query(state["request"])
    return {"policies": retrieve_policies(query)}


def analyze_node(state: AgentState) -> AgentState:
    gate_status = deterministic_gate(state["request"])
    draft = generate_draft(state["request"], state["policies"], gate_status)
    return {"gate_status": gate_status, "draft": draft.model_dump()}


def route_after_analysis(state: AgentState) -> Literal["review", "finish"]:
    return "finish" if state["gate_status"] == "BLOCKED" else "review"


def review_node(state: AgentState) -> AgentState:
    approved = interrupt(
        {
            "question": "Approve this simulated supplier record?",
            "request_id": state["request"]["request_id"],
            "gate_status": state["gate_status"],
            "draft": state["draft"],
        }
    )
    return {"human_approved": bool(approved)}


def route_after_review(state: AgentState) -> Literal["record", "finish"]:
    return "record" if state["human_approved"] else "finish"


@traceable(run_type="tool", name="Record approved supplier")
def record_node(state: AgentState) -> AgentState:
    ACTION_LOG.parent.mkdir(parents=True, exist_ok=True)
    request_id = str(state["request"]["request_id"])

    if ACTION_LOG.exists():
        for line in ACTION_LOG.read_text().splitlines():
            if line and json.loads(line).get("request_id") == request_id:
                return {
                    "final_status": "APPROVED_ALREADY_RECORDED",
                    "action_record": str(ACTION_LOG),
                }

    record = {
        "request_id": request_id,
        "supplier_name": state["request"]["supplier_name"],
        "status": "APPROVED",
        "gate_status": state["gate_status"],
        "policy_ids": [
            policy["policy_id"] for policy in state["policies"]
        ],
        "approved_at": datetime.now(timezone.utc).isoformat(),
    }
    with ACTION_LOG.open("a") as handle:
        handle.write(json.dumps(record) + "\n")
    return {
        "final_status": "APPROVED_AND_RECORDED",
        "action_record": str(ACTION_LOG),
    }


def finish_node(state: AgentState) -> AgentState:
    if state["gate_status"] == "BLOCKED":
        return {"final_status": "BLOCKED"}
    return {"final_status": "REJECTED_BY_HUMAN"}


def build_graph():
    builder = StateGraph(AgentState)
    builder.add_node("retrieve", retrieve_node)
    builder.add_node("analyze", analyze_node)
    builder.add_node("review", review_node)
    builder.add_node("record", record_node)
    builder.add_node("finish", finish_node)
    builder.add_edge(START, "retrieve")
    builder.add_edge("retrieve", "analyze")
    builder.add_conditional_edges(
        "analyze",
        route_after_analysis,
        {"review": "review", "finish": "finish"},
    )
    builder.add_conditional_edges(
        "review",
        route_after_review,
        {"record": "record", "finish": "finish"},
    )
    builder.add_edge("record", END)
    builder.add_edge("finish", END)
    return builder.compile(checkpointer=InMemorySaver())


GRAPH = build_graph()


@traceable(run_type="chain", name="Governed supplier approval agent")
def run_agent(
    request: dict[str, object], auto_approve: bool | None = None
) -> AgentState:
    config = {"configurable": {"thread_id": str(request["request_id"])}}
    result = GRAPH.invoke({"request": request}, config=config)
    pending = result.get("__interrupt__", ())

    if pending:
        print("\nHUMAN REVIEW REQUIRED")
        print(json.dumps(pending[0].value, indent=2))
        approved = auto_approve
        if approved is None:
            approved = input("Type approve to write the record: ").strip().lower() == "approve"
        result = GRAPH.invoke(Command(resume=approved), config=config)

    return result


def main() -> None:
    request = json.loads(Path("data/sample_request.json").read_text())
    result = run_agent(request)
    print("\nFINAL STATUS:", result["final_status"])
    print("DRAFT:", json.dumps(result["draft"], indent=2))
    if result.get("action_record"):
        print("ACTION RECORD:", result["action_record"])


if __name__ == "__main__":
    main()

How does this code stay governed?

  • The DecisionDraft model requires Gemini to return a recommendation, rationale, and policy citations in a predictable structure.
  • The deterministic_gate() function blocks unclear sanctions screening and incomplete customer-data security reviews.
  • The review_node() function pauses permitted requests before any approval record can be written.
  • The record_node() function checks the request ID before appending a record, which makes the write idempotent.
  • Save agent.py.
  • Confirm that the saved file sits beside retrieval.py in your project folder.

Can't save or import the agent?

Check that the filename is exactly agent.py. Confirm that retrieval.py is in the same folder.

Keep the existing virtual environment active so the installed packages remain available. Help me diagnose an agent import problem

Follow the guarded route

The graph separates probabilistic reasoning from deterministic authorization. Each node has one responsibility, so the approval boundary stays visible in code and in traces.

  • Search agent.py for deterministic_gate().
  • Confirm that a non-clear sanctions result returns BLOCKED.
  • Confirm that a missing security review for a customer-data supplier returns BLOCKED.
  • Search agent.py for review_node().
  • Confirm that interrupt() receives the request ID, gate status, and structured draft.
  • Search agent.py for record_node().
  • Confirm that the request ID check occurs before the append operation.

Why keep the write in a separate node?

An interrupted LangGraph node starts again when the graph resumes. Placing the write after the interrupt prevents the approval side effect from running during that restart.

The request ID check adds a second layer of protection. Repeated approval attempts return an existing-record status instead of adding duplicate lines.

Approve the supplier and verify the trace

The sample supplier passes the sanctions and security gates. Its annual spend still requires director review, so the graph must pause before creating a record.

Before you run the workflow, which gate status do you expect for a supplier requesting USD 75,000 in annual spend?

  • Start the governed workflow by running this command:
python agent.py

What happens during this run?

The graph converts the supplier request into a retrieval query. It passes the retrieved policies and deterministic gate status to Gemini for a structured draft.

The graph reaches interrupt() before record_node(). The terminal remains at the human review checkpoint until you provide a decision.

You will see HUMAN REVIEW REQUIRED followed by an approval payload. The payload includes SUP-2026-001 and DIRECTOR_REVIEW_REQUIRED.

This approval writes only fictional supplier data to a local ignored file. You remain in control of the simulated side effect.

  • Type approve at the terminal prompt.
  • Press Enter to resume the graph.

You will see FINAL STATUS: APPROVED_AND_RECORDED and the action-record path artifacts/approved_suppliers.jsonl.

Does the workflow stop before approval?

Check that data/sample_request.json still contains a clear sanctions result and a completed security review. A blocked request correctly skips the human approval interrupt.

Confirm that your Gemini credentials remain in .env if the run stops during analysis. Help me troubleshoot the governed workflow

The first approval worked. Now test whether the write remains safe when the same request is approved again.

Before you repeat the run, do you expect another JSON line or an already-recorded status?

  • Repeat the governed workflow by running the same command:
python agent.py

What does the repeat run test?

The workflow processes the same request ID again. After approval, record_node() searches the existing JSONL records before attempting another append.

  • Type approve at the terminal prompt.
  • Press Enter to resume the graph.

You will see FINAL STATUS: APPROVED_ALREADY_RECORDED. This confirms that the repeated approval did not create a duplicate record.

  • Expand the artifacts folder in your project file tree.
  • Open approved_suppliers.jsonl.

You will see one JSON line for SUP-2026-001. It includes the approved status, director-review gate, retrieved policy IDs, and a UTC approval timestamp.

  • Return to your LangSmith account from the setup step.
  • Select the governed-supplier-agent project.
  • Open the trace generated by your latest workflow run.

You will see nested retrieval, embedding, Gemini, and record runs. The trace shows where evidence entered the workflow and where code controlled the write.

Your supplier agent now grounds its recommendation, enforces deterministic rules, pauses for human judgment, and avoids duplicate writes. Next, you will test those guarantees across several risk cases.

Evaluate and Trace the Workflow

Your governed agent now keeps the model away from the write boundary. A single successful approval still cannot prove the workflow stays safe across different risk conditions.

Regression testing checks the same expectations against three distinct requests. It catches changes in retrieval, routing, or citations before they reach a demo.

LangSmith traces expose each nested retrieval call. They also show the Gemini run behind every result.

In this step, get ready to:
  • Define three supplier-risk cases with expected gate statuses and policy citations.
  • Run the evaluator to check retrieval, deterministic routing, and citation validity.
  • Inspect the corresponding retrieval and Gemini runs in LangSmith.
Define three supplier-risk cases

A regression case gives one supplier request a known expected outcome. These cases cover high spend, missing security review, and unclear sanctions screening.

  • In your project editor from earlier, create evaluate.py beside agent.py.
  • Add the imports and the high-spend case by copying this code into evaluate.py:
import json

from agent import deterministic_gate, generate_draft
from retrieval import build_request_query, retrieve_policies

CASES = [
    {
        "name": "high_spend",
        "request": {
            "request_id": "EVAL-001",
            "supplier_name": "High Spend Labs",
            "annual_spend_usd": 75000,
            "handles_customer_data": False,
            "security_review_complete": False,
            "sanctions_screening": "clear",
        },
        "expected_gate": "DIRECTOR_REVIEW_REQUIRED",
        "expected_policy_id": "POL-FIN-002",
    },
]

What does this code define?

  • The imports reuse the production functions from agent.py and retrieval.py.
  • The high_spend request represents a clear supplier whose annual spend requires director review.
  • The expected policy is POL-FIN-002 because that policy governs high-spend approval.
  • Save evaluate.py.
  • Confirm that evaluate.py appears beside agent.py in your editor's file sidebar.

Is the new file missing?

  • Check that you created evaluate.py inside the same project directory as agent.py.
  • Check that the filename ends with .py.

Help me find or create evaluate.py in the correct project directory.

The remaining cases target the two deterministic blocking conditions. Their expected policy IDs let the evaluator check whether retrieval surfaced the evidence that explains each block.

  • In evaluate.py, locate the closing bracket for CASES.
  • Insert these two case dictionaries immediately above that closing bracket:
    {
        "name": "missing_security_review",
        "request": {
            "request_id": "EVAL-002",
            "supplier_name": "Customer Data Systems",
            "annual_spend_usd": 10000,
            "handles_customer_data": True,
            "security_review_complete": False,
            "sanctions_screening": "clear",
        },
        "expected_gate": "BLOCKED",
        "expected_policy_id": "POL-SEC-001",
    },
    {
        "name": "unclear_sanctions",
        "request": {
            "request_id": "EVAL-003",
            "supplier_name": "Global Components",
            "annual_spend_usd": 8000,
            "handles_customer_data": False,
            "security_review_complete": False,
            "sanctions_screening": "pending",
        },
        "expected_gate": "BLOCKED",
        "expected_policy_id": "POL-CMP-003",
    },

What do these cases test?

  • The missing_security_review case must retrieve POL-SEC-001 and finish with BLOCKED.
  • The unclear_sanctions case must retrieve POL-CMP-003 and finish with BLOCKED.
  • Both requests exercise code-owned gates without using the human approval path.
  • Save evaluate.py.
  • Use your editor's search box to search for "name":.
  • Confirm that the search returns three matches.

Do you see fewer than three cases?

  • Check that both dictionaries sit inside the opening and closing brackets for CASES.
  • Check that a comma separates each dictionary.
  • Check that the three case names are high_spend, missing_security_review, and unclear_sanctions.

Help me fix the CASES list in evaluate.py.

Build and run the evaluator

The evaluator checks more than the final gate status. It also verifies that the expected policy was retrieved and that every model citation came from the retrieved evidence.

  • In evaluate.py, place the evaluator function below the closing bracket for CASES by copying this code:
def main() -> None:
    passed = 0
    for case in CASES:
        request = case["request"]
        policies = retrieve_policies(build_request_query(request))
        gate = deterministic_gate(request)
        draft = generate_draft(request, policies, gate)
        retrieved_ids = {str(policy["policy_id"]) for policy in policies}
        cited_ids = set(draft.cited_policy_ids)
        case_passed = (
            gate == case["expected_gate"]
            and case["expected_policy_id"] in retrieved_ids
            and case["expected_policy_id"] in cited_ids
            and cited_ids.issubset(retrieved_ids)
        )
        passed += int(case_passed)
        print(
            json.dumps(
                {
                    "case": case["name"],
                    "passed": case_passed,
                    "gate": gate,
                    "retrieved_ids": sorted(retrieved_ids),
                    "cited_ids": sorted(cited_ids),
                }
            )
        )

    print(f"EVALUATION: {passed}/{len(CASES)} passed")

How does the evaluation work?

  • Each request passes through build_request_query() and retrieve_policies() to collect policy evidence.
  • The deterministic_gate() result must equal the case's expected gate status.
  • The expected policy ID must appear in both the retrieved IDs and the model's cited IDs.
  • The issubset() check prevents the model from citing a policy that retrieval did not supply.
  • Save evaluate.py.
  • Use your editor's search box to search for def main().
  • Confirm that the search returns one evaluator function.

Does the evaluator function look incomplete?

  • Check that case_passed contains four conditions.
  • Check that the final print() sits outside the for loop.
  • Check that every line inside main() uses consistent indentation.

Help me check the structure of main() in evaluate.py.

  • Add the script entry point at the bottom of evaluate.py by copying this code:
if __name__ == "__main__":
    main()

Why add an entry point?

The entry point calls main() when you run evaluate.py as a script. Importing the file from another module does not start the evaluation automatically.

  • Save evaluate.py.

Before you run the evaluator, predict whether all three requests will satisfy the retrieval, routing, and citation checks.

  • Run all three regression cases by running this command:
python evaluate.py

What does this command test?

The command sends three fictional requests through policy retrieval, deterministic routing, and grounded Gemini analysis. Each printed result summarizes the evidence and decision checks for one case.

You should see one JSON object per case. The final line should read EVALUATION: 3/3 passed.

Did one of the cases fail?

  • Check the failed case's gate value against its expected_gate value.
  • Compare retrieved_ids with the case's expected policy ID.
  • Compare cited_ids with the policies supplied to Gemini.

Help me investigate a failed evaluation case.

Treat failures as quality signals

A failed assertion points to a retrieval, routing, or citation problem. Compare the trace with the expected case before changing any code.

Keep deterministic_gate() aligned with the business rules. A model response must never redefine those controls.

✔️ Awesome, I've got everything!

Great. Your three-case evaluator is saved and producing a complete pass report.

ⓧ I'd like to double check the full code

import json

from agent import deterministic_gate, generate_draft
from retrieval import build_request_query, retrieve_policies

CASES = [
    {
        "name": "high_spend",
        "request": {
            "request_id": "EVAL-001",
            "supplier_name": "High Spend Labs",
            "annual_spend_usd": 75000,
            "handles_customer_data": False,
            "security_review_complete": False,
            "sanctions_screening": "clear",
        },
        "expected_gate": "DIRECTOR_REVIEW_REQUIRED",
        "expected_policy_id": "POL-FIN-002",
    },
    {
        "name": "missing_security_review",
        "request": {
            "request_id": "EVAL-002",
            "supplier_name": "Customer Data Systems",
            "annual_spend_usd": 10000,
            "handles_customer_data": True,
            "security_review_complete": False,
            "sanctions_screening": "clear",
        },
        "expected_gate": "BLOCKED",
        "expected_policy_id": "POL-SEC-001",
    },
    {
        "name": "unclear_sanctions",
        "request": {
            "request_id": "EVAL-003",
            "supplier_name": "Global Components",
            "annual_spend_usd": 8000,
            "handles_customer_data": False,
            "security_review_complete": False,
            "sanctions_screening": "pending",
        },
        "expected_gate": "BLOCKED",
        "expected_policy_id": "POL-CMP-003",
    },
]


def main() -> None:
    passed = 0
    for case in CASES:
        request = case["request"]
        policies = retrieve_policies(build_request_query(request))
        gate = deterministic_gate(request)
        draft = generate_draft(request, policies, gate)
        retrieved_ids = {str(policy["policy_id"]) for policy in policies}
        cited_ids = set(draft.cited_policy_ids)
        case_passed = (
            gate == case["expected_gate"]
            and case["expected_policy_id"] in retrieved_ids
            and case["expected_policy_id"] in cited_ids
            and cited_ids.issubset(retrieved_ids)
        )
        passed += int(case_passed)
        print(
            json.dumps(
                {
                    "case": case["name"],
                    "passed": case_passed,
                    "gate": gate,
                    "retrieved_ids": sorted(retrieved_ids),
                    "cited_ids": sorted(cited_ids),
                }
            )
        )

    print(f"EVALUATION: {passed}/{len(CASES)} passed")


if __name__ == "__main__":
    main()
Inspect the nested LangSmith runs

A passing report shows that the assertions held. The traces reveal how retrieval and model reasoning produced those results.

Before you inspect the recent runs, predict which inputs and outputs should connect each supplier case to its policy evidence.

  • Return to your LangSmith account from earlier.
  • Open the governed-supplier-agent project.
  • Find the recent runs created when you executed evaluate.py.
  • Open a Search supplier policies run.
  • Expand its nested embedding run.
  • Confirm that the retrieved output contains policy IDs for the matching supplier case.
  • Open the corresponding Gemini grounded supplier recommendation run.
  • Confirm that its structured output cites only policy IDs from the retrieved evidence.

You should see recent retrieval and Gemini runs for the evaluation calls. Their inputs and outputs give you a traceable path from each fictional request to its final case result.

Cannot find the evaluation runs?

  • Confirm that you are viewing the governed-supplier-agent project.
  • Confirm that tracing remains enabled in the .env file from earlier.
  • Run the evaluator again after confirming the tracing configuration.

Help me find my recent evaluation traces in LangSmith.

That is the production-minded proof complete. Your agent now has a repeatable quality check and an inspectable trail behind its recommendations.

Secret mission

Expose Policy Search as an MCP Tool

Your validated policy retriever currently serves one governed agent. Expose it as a typed, read-only MCP tool so other clients can retrieve evidence without gaining access to the approval-record write.

Clean Up Your Resources

Clean Up Your Resources

This project has no ongoing cost while the Gemini Developer API stays on its Free tier with LangSmith kept payment-free. Choose whether to keep the resources available, pause tracing for later, or delete the generated state.

Resources you used:

  • Gemini API key created in Google AI Studio.
  • LangSmith API key for the Developer account.
  • LangSmith project governed-supplier-agent containing the workflow traces.
  • Local LanceDB table at data/lancedb/supplier_policies.
  • Any idempotent approval record stored at artifacts/approved_suppliers.jsonl.

Keep everything running

No action needed. Choose this if you want to keep testing the governed agent or the read-only MCP tool.

  • Keep data/lancedb/ so future policy searches can reuse the embedded table.
  • Keep artifacts/ so the idempotent write check can recognize an existing approval record.
  • Keep the governed-supplier-agent project so you can compare future traces with the evaluation runs.
  • Keep the LangSmith account payment-free so its monthly trace limit remains a hard usage boundary.

The terminal agent and MCP server stop when their processes exit. Keeping these resources does not leave a local service running.

Pause - I'll come back to this later

Pause LangSmith tracing while preserving the local policy table and approval record for another session.

  • Return to .env in the project directory from earlier.
  • Find LANGSMITH_TRACING=true.
  • Replace it with LANGSMITH_TRACING=false.
  • Save .env.

Future runs keep working locally without adding LangSmith traces. Your Gemini key and LangSmith key remain available for when you return.

Delete - I don't want to use this again

Deleting these resources is permanent. Your source files remain available if you want to rebuild the generated state later.

Revoke external access:

  • Sign in to Google AI Studio.
  • Return to the API-key page you used in Step 1.
  • Revoke the Gemini API key created for this project.
  • Sign in to LangSmith.
  • Revoke the project API key from your account settings.
  • Open the governed-supplier-agent project.
  • Delete the project from its settings to remove its stored traces.

The revoked keys can no longer authenticate future project runs. The deleted LangSmith project no longer retains its trace history.

Remove local generated state:

  • Remove the local vector database and approval records from the project directory from earlier by running:
rm -rf data/lancedb/ artifacts/

What Does This Command Remove?

The command removes data/lancedb/ with its embedded supplier-policy table. It also removes artifacts/ with any simulated approval record.

  • Confirm in your Linux file manager that data/lancedb/ no longer appears inside data/.
  • Confirm in the project directory that artifacts/ no longer appears.

Cleanup complete. Your generated policy index and simulated approval record are gone.

Nice Work!

Nice Work!

Mission complete! You built a governed supplier-approval agent that retrieves policy evidence before Gemini drafts a recommendation. Deterministic controls keep the final write behind human approval.

You've learned how to:

  • Built a structured supplier workflow with Gemini that produces schema-validated recommendations.
  • Created a local RAG pipeline backed by LanceDB. Each recommendation now carries retrieved policy evidence.
  • Kept deterministic controls in charge of gate status. Added a LangGraph interrupt before the idempotent approval write. Validated policy retrieval with a three-case evaluation suite. The same suite checked gate routing. It also checked citation validity. Inspected the resulting traces in LangSmith.
  • Completed the Secret Mission by exposing policy retrieval through a typed read-only MCP tool. Proved the tool contract with an in-memory client.

Ready to quiz yourself?