Route AI Requests with LiteLLM

Deploy a LiteLLM gateway that routes Gemini requests with Jev on Google Cloud.

Introduction

30 Second Summary

A one-word question can be sent to the same powerful AI as a complex design problem. That makes routine work more expensive than it needs to be.

In this project, you will deploy a cost-aware AI gateway using LiteLLM version 1.104.0 on Google Cloud Run. TypeSafe AI Jev will classify each request so the gateway can choose a suitable Gemini model.

What You'll Build

Picture sending prompts through your live gateway while the Admin UI proves that simple work reached the cheaper model.

By the end of this project, you'll have:

  • A live gateway you can call from a benchmark script to receive authenticated AI responses.
  • Repeatable baseline and routed benchmarks with CSV evidence covering model choice plus latency plus classifier cost.
  • An Admin UI evidence trail showing each Jev decision beside spend attributed to your virtual key.
  • Secret Mission: A tiny-budget virtual key that proves LiteLLM blocks further requests with HTTP 422.

Are there any prerequisites?

You need a billed Google Cloud project with permission to create project services plus IAM bindings. You also need a TypeSafe AI account for the Jev key.

Before We Start

This step is your moment to commit to the gateway's purpose before any hands-on work. You're building a public, cost-aware LiteLLM AI gateway that routes OpenAI-compatible requests by Jev complexity tier to reserve the stronger model for prompts that need it.

Set Up Cloud Shell and Credentials

Your cost-aware gateway needs an authenticated workspace for the cloud work ahead. Google Cloud Shell provides that workspace in your browser.

Cloud Shell keeps macOS setup out of the critical path. It also includes the browser editor and Docker for the deployment workflow.

You will confirm the billed project before pinning Terraform to 1.16.5. A verified TypeSafe AI credential will prepare Jev for the routing step.

In this step, get ready to:
  • Open an authenticated Cloud Shell workspace for the billed project.
  • Confirm Terraform 1.16.5 is available on PATH.
  • Verify a TypeSafe API key against the available Jev models.
Open Cloud Shell and verify the project

Cloud Shell receives the active project selected in Google Cloud. Checking that project now prevents every later command from targeting the wrong account.

  • Open shell.cloud.google.com in your browser.
  • Use the project selector in the page header to choose your billed Google Cloud project.
  • Wait for the Cloud Shell terminal prompt to finish loading.
  • Launch Cloud Shell Editor from the terminal by running this command:
cloudshell edit .

What does this command do?

The cloudshell edit . command opens the current Cloud Shell directory in Cloud Shell Editor. The terminal remains available beside the editor.

You should see the Cloud Shell Editor with a file explorer and an active terminal. No project files exist yet.

  • Confirm which Google Cloud project is active by running this command:
gcloud config get-value project

What does this command confirm?

The gcloud configuration controls which project receives later infrastructure commands. This check prints the current project ID without creating any resources.

You should see the ID of your billed Google Cloud project.

Seeing the wrong project ID?

  • Return to the project selector in the page header.
  • Choose the billed project you plan to use for the gateway.
  • Restart the Cloud Shell session from the selected project.
  • Rerun the project check after the terminal loads.

Ask for help if the project still does not match. Help me connect Cloud Shell to my billed Google Cloud project

  • Record the verified project ID here for the next step: your-project-id.
Verify Terraform and Docker

Cloud Shell images can change over time. A version check ensures your infrastructure configuration uses the exact Terraform release expected by this project.

  • Check the Terraform version available in Cloud Shell by running this command:
terraform version

What does this command show?

The command prints the active Terraform CLI version. This project requires exactly Terraform v1.16.5 so its configuration matches the pinned release.

✔️ I see version 1.16.5

The required Terraform release is already on PATH. Your Cloud Shell session is ready for the infrastructure files in the next step.

ⓧ I see an older version

The older binary could produce different validation behavior. Installing the pinned Linux AMD64 binary in $HOME/bin gives this project a predictable Terraform release.

  • Install the checksum-verified Terraform binary in your persistent home directory by running these commands:
curl --remote-name https://releases.hashicorp.com/terraform/1.16.5/terraform_1.16.5_linux_amd64.zip && \
echo "2bc2fcfff033265c9e02ca0351f01794eb122f62a9b2a49a3294b9e49eaab5e4  terraform_1.16.5_linux_amd64.zip" | sha256sum -c && \
mkdir -p "$HOME/bin" && \
unzip -o terraform_1.16.5_linux_amd64.zip -d "$HOME/bin" && \
echo 'export PATH="$HOME/bin:$PATH"' >> "$HOME/.bashrc" && \
export PATH="$HOME/bin:$PATH" && \
rm terraform_1.16.5_linux_amd64.zip && \
terraform version

How does this install stay safe?

  • The first command downloads the official Terraform 1.16.5 Linux AMD64 archive.
  • The sha256sum -c check compares the download with HashiCorp's published checksum.
  • The archive is extracted into $HOME/bin so it persists in your Cloud Shell home directory.
  • The updated PATH places the pinned binary ahead of any older installation.

You should see the checksum report OK. The final output should show Terraform v1.16.5.

Checksum or version does not match?

  • Stop if the checksum check does not report OK.
  • Rerun the installation block to download a fresh copy from HashiCorp.
  • Confirm the final version output comes from $HOME/bin.

Use this if the pinned binary still does not run. Help me troubleshoot the Terraform 1.16.5 installation in Cloud Shell

ⓧ Command not found

Cloud Shell can run without a preinstalled Terraform CLI. The persistent home directory provides a suitable location for the pinned binary.

  • Install the checksum-verified Terraform binary in your persistent home directory by running these commands:
curl --remote-name https://releases.hashicorp.com/terraform/1.16.5/terraform_1.16.5_linux_amd64.zip && \
echo "2bc2fcfff033265c9e02ca0351f01794eb122f62a9b2a49a3294b9e49eaab5e4  terraform_1.16.5_linux_amd64.zip" | sha256sum -c && \
mkdir -p "$HOME/bin" && \
unzip -o terraform_1.16.5_linux_amd64.zip -d "$HOME/bin" && \
echo 'export PATH="$HOME/bin:$PATH"' >> "$HOME/.bashrc" && \
export PATH="$HOME/bin:$PATH" && \
rm terraform_1.16.5_linux_amd64.zip && \
terraform version

How does this install stay safe?

  • The first command downloads the official Terraform 1.16.5 Linux AMD64 archive.
  • The sha256sum -c check compares the download with HashiCorp's published checksum.
  • The archive is extracted into $HOME/bin so it persists in your Cloud Shell home directory.
  • The updated PATH makes the new binary available to the current shell.

You should see the checksum report OK. The final output should show Terraform v1.16.5.

Terraform still unavailable?

  • Stop if the checksum check does not report OK.
  • Rerun the installation block to download a fresh archive.
  • Confirm that $HOME/bin appears at the start of PATH.

Use this if the command remains unavailable. Help me fix Terraform command not found in Cloud Shell

  • Confirm Docker is available in Cloud Shell by running this command:
docker --version

Why check Docker now?

Docker builds the LiteLLM image that Cloud Run deploys later. This check confirms the CLI is available before any project files are created.

You should see Docker version information in the terminal.

Docker command unavailable?

Restart the Cloud Shell session to load its current tool image. Rerun the version check after the terminal prompt returns.

Use this if Docker remains unavailable. Help me restore Docker in Google Cloud Shell

Create and verify the TypeSafe API key

The next value is a live credential. Keeping it only in the shell environment prevents it from appearing in project files or terminal output.

  • Open the TypeSafe API keys page in a new browser tab.
  • Use the key creation control on the page to create an API key.

The page presents the generated key for you to save. Treat this value like a password.

  • Copy the generated key from the TypeSafe page.
  • Switch back to the Cloud Shell terminal from earlier.
  • Load the key through a hidden terminal prompt by running this command:
read -s TYPESAFE_API_KEY; export TYPESAFE_API_KEY; echo

How does the hidden prompt protect the key?

  • The read -s TYPESAFE_API_KEY command stores your pasted value without displaying its characters.
  • The export TYPESAFE_API_KEY command makes the credential available to later processes in this shell.
  • The final echo returns the prompt to a clean line without printing the credential.

The blank line confirms that the hidden input finished without exposing your key.

Before you run the final checks, do you expect all three to confirm the same prepared environment?

  • Verify the active project, Terraform release, and TypeSafe credential by running these commands:
gcloud config get-value project
terraform version
curl https://api.typesafe.ai/v1/models -H "Authorization: Bearer $TYPESAFE_API_KEY"

What do these checks prove?

  • The first check identifies the Google Cloud project that receives later deployment commands.
  • The second check confirms the Terraform CLI matches the project's pinned release.
  • The final request uses the environment-held credential to retrieve the models available to your TypeSafe account.

You should see your billed project ID first. The Terraform output should show Terraform v1.16.5.

The JSON response should list at least one available Jev model. This may include jev-latest.

Your authenticated workspace is ready for deployment.

TypeSafe request rejected?

  • Create a fresh key from the TypeSafe API keys page if the copied value is incomplete.
  • Rerun the hidden input command to replace the current shell value.
  • Rerun the models request after the new value is loaded.

Use this if the models request still fails. Help me troubleshoot my TypeSafe models request

Setup Checkpoint

Cloud Shell is authenticated to the billed Google Cloud project. Cloud Shell Editor is available.

Terraform version 1.16.5 is available on PATH.

The TypeSafe API key is verified through https://api.typesafe.ai/v1/models. It is currently held only in the shell environment variable TYPESAFE_API_KEY.

Your cloud workspace now has the authenticated tools and credential required by the gateway. Next, you will create the project files and deploy the Admin UI.

Deploy the Gateway and Admin UI

Your authenticated Google Cloud Shell session can now turn the gateway design into working cloud infrastructure. Your verified TypeSafe credential is also ready for secure storage.

The fastest useful milestone is a live LiteLLM Admin UI backed by durable data. In this step, Terraform creates the supporting services before Cloud Run launches the gateway.

In this step, get ready to:
  • Create the gateway files and validate the Terraform configuration.
  • Store the TypeSafe credential and push the LiteLLM image.
  • Deploy the gateway and verify its Admin UI.
Create and validate the project

The project combines container configuration with infrastructure as code. The files define the same stack that Terraform later creates in your billed project.

  • Switch back to the Cloud Shell terminal from earlier.
  • Create the `jev-litellm-gateway` folder in your Cloud Shell home directory by running these commands:
mkdir jev-litellm-gateway
cd jev-litellm-gateway

What do these commands do?

The first command creates the project folder. The second command makes that folder the active terminal location.

  • Create the project files inside `jev-litellm-gateway` by running:
touch .gitignore Dockerfile benchmark.py config.yaml main.tf outputs.tf terraform.tfvars.example variables.tf versions.tf

What does this command do?

The command creates the nine files that hold the gateway image definition, LiteLLM configuration, benchmark, and Terraform stack.

  • Create the local Terraform values file from the example by running:
cp terraform.tfvars.example terraform.tfvars

Why use a separate values file?

The example documents the required setting. The local `terraform.tfvars` file stores your project-specific value and stays outside version control.

  • Confirm that the files exist by running:
ls

What should you see?

You should see `Dockerfile`, `benchmark.py`, `config.yaml`, the Terraform files, and `terraform.tfvars` in the output. The hidden `.gitignore` file does not appear in a standard listing.

  • Return to the Cloud Shell Editor from earlier.
  • Refresh its file explorer so the `jev-litellm-gateway` folder appears.
  • Expand the `jev-litellm-gateway` folder.
  • Use the second tab below to paste the matching contents into every project file.

✔️ Awesome, I've got everything!

Great. Save every file before continuing to the Terraform checks.

ⓧ I'd like to double check the full code

Compare each file in `jev-litellm-gateway` with the complete versions below. Keep `terraform.tfvars` as a local copy of `terraform.tfvars.example` with your billed project ID.

.terraform/
*.tfstate
*.tfstate.*
terraform.tfvars
benchmark-*.csv
.typesafe-key

What does this file protect?

The ignore rules keep Terraform state, local values, benchmark output, and temporary TypeSafe key files out of version control.

FROM docker.litellm.ai/berriai/litellm:1.104.0
WORKDIR /app
COPY config.yaml /app/config.yaml
EXPOSE 8080/tcp
CMD ["--port", "8080", "--config", "/app/config.yaml"]

What does this file build?

The image pins LiteLLM to `1.104.0`. It copies the gateway configuration into the image and starts the proxy on port `8080`.

model_list:
  - model_name: gemini-flash-lite
    litellm_params:
      model: vertex_ai/gemini-3.5-flash-lite
      vertex_project: os.environ/VERTEXAI_PROJECT
      vertex_location: os.environ/VERTEXAI_LOCATION
  - model_name: gemini-flash
    litellm_params:
      model: vertex_ai/gemini-3.5-flash
      vertex_project: os.environ/VERTEXAI_PROJECT
      vertex_location: os.environ/VERTEXAI_LOCATION

general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
  database_url: os.environ/DATABASE_URL

What does this configuration define?

The two aliases connect LiteLLM to Gemini 3.5 Flash-Lite and Gemini 3.5 Flash through Vertex AI. The general settings connect authentication and persistent data to environment-provided secrets.

terraform {
  required_version = "= 1.16.5"
  required_providers {
    google = {
      source  = "hashicorp/google"
      version = "8.5.0"
    }
    random = {
      source  = "hashicorp/random"
      version = "3.9.1"
    }
  }
}
provider "google" {
  project = var.project_id
  region  = var.region
}

Why pin these versions?

The file requires Terraform `1.16.5`. It also fixes the Google provider at `8.5.0` and the Random provider at `3.9.1` so the configuration uses known resource schemas.

variable "project_id" {
  description = "The billed Google Cloud project ID."
  type        = string
}
variable "region" {
  description = "The region for Cloud Run, Cloud SQL, and Artifact Registry."
  type        = string
  default     = "us-central1"
}
variable "vertex_location" {
  description = "The Vertex AI location used by both Gemini deployments."
  type        = string
  default     = "global"
}

What do these variables control?

The variables select your billed project. They also place the cloud resources in `us-central1` while Gemini requests use the `global` location.

locals {
  required_services = toset([
    "aiplatform.googleapis.com",
    "artifactregistry.googleapis.com",
    "iam.googleapis.com",
    "run.googleapis.com",
    "secretmanager.googleapis.com",
    "sqladmin.googleapis.com",
  ])
  repository_id  = "jev-litellm"
  service_name   = "jev-litellm-gateway"
  image_uri      = "${var.region}-docker.pkg.dev/${var.project_id}/${local.repository_id}/litellm:1.104.0"
  runtime_member = "serviceAccount:${google_service_account.gateway.email}"
  database_url   = "postgresql://litellm:${random_password.database.result}@localhost/litellm?host=/cloudsql/${google_sql_database_instance.postgres.connection_name}&connection_limit=2"
}

resource "google_project_service" "apis" {
  for_each = local.required_services
  project = var.project_id
  service = each.value
  disable_on_destroy = false
}

resource "google_artifact_registry_repository" "images" {
  project = var.project_id
  location = var.region
  repository_id = local.repository_id
  description = "LiteLLM gateway images"
  format = "DOCKER"
  depends_on = [google_project_service.apis]
}

resource "google_service_account" "gateway" {
  project = var.project_id
  account_id = "jev-litellm"
  display_name = "Jev LiteLLM gateway"
  depends_on = [google_project_service.apis]
}

resource "google_project_iam_member" "runtime_roles" {
  for_each = toset(["roles/aiplatform.user", "roles/cloudsql.client"])
  project = var.project_id
  role = each.value
  member = local.runtime_member
}

resource "random_password" "database" {
  length = 32
  special = false
}
resource "random_password" "master" {
  length = 48
  special = false
}
resource "random_password" "salt" {
  length = 48
  special = false
}

resource "google_sql_database_instance" "postgres" {
  project = var.project_id
  name = "jev-litellm-postgres"
  region = var.region
  database_version = "POSTGRES_15"
  settings {
    tier = "db-f1-micro"
    edition = "ENTERPRISE"
    disk_type = "PD_SSD"
    disk_size = 10
    disk_autoresize = false
    connector_enforcement = "REQUIRED"
    backup_configuration { enabled = false }
    ip_configuration { ipv4_enabled = true }
  }
  deletion_protection = false
  depends_on = [google_project_service.apis]
}

resource "google_sql_database" "litellm" {
  project = var.project_id
  name = "litellm"
  instance = google_sql_database_instance.postgres.name
}
resource "google_sql_user" "litellm" {
  project = var.project_id
  name = "litellm"
  instance = google_sql_database_instance.postgres.name
  password = random_password.database.result
}

resource "google_secret_manager_secret" "master_key" {
  project = var.project_id
  secret_id = "jev-litellm-master-key"
  replication { auto {} }
  depends_on = [google_project_service.apis]
}
resource "google_secret_manager_secret" "salt_key" {
  project = var.project_id
  secret_id = "jev-litellm-salt-key"
  replication { auto {} }
  depends_on = [google_project_service.apis]
}
resource "google_secret_manager_secret" "database_url" {
  project = var.project_id
  secret_id = "jev-litellm-database-url"
  replication { auto {} }
  depends_on = [google_project_service.apis]
}
resource "google_secret_manager_secret" "typesafe_api_key" {
  project = var.project_id
  secret_id = "jev-litellm-typesafe-api-key"
  replication { auto {} }
  depends_on = [google_project_service.apis]
}

resource "google_secret_manager_secret_version" "master_key" {
  secret = google_secret_manager_secret.master_key.id
  secret_data = "sk-${random_password.master.result}"
}
resource "google_secret_manager_secret_version" "salt_key" {
  secret = google_secret_manager_secret.salt_key.id
  secret_data = "sk-${random_password.salt.result}"
}
resource "google_secret_manager_secret_version" "database_url" {
  secret = google_secret_manager_secret.database_url.id
  secret_data = local.database_url
  depends_on = [google_sql_database.litellm, google_sql_user.litellm]
}

locals {
  runtime_secrets = {
    master_key = google_secret_manager_secret.master_key.id
    salt_key = google_secret_manager_secret.salt_key.id
    database_url = google_secret_manager_secret.database_url.id
    typesafe_api_key = google_secret_manager_secret.typesafe_api_key.id
  }
}

resource "google_secret_manager_secret_iam_member" "runtime_secret_access" {
  for_each = local.runtime_secrets
  project = var.project_id
  secret_id = each.value
  role = "roles/secretmanager.secretAccessor"
  member = local.runtime_member
}

resource "google_cloud_run_v2_service" "gateway" {
  project = var.project_id
  name = local.service_name
  location = var.region
  ingress = "INGRESS_TRAFFIC_ALL"
  deletion_protection = false
  template {
    service_account = google_service_account.gateway.email
    timeout = "300s"
    scaling {
      min_instance_count = 0
      max_instance_count = 1
    }
    containers {
      image = local.image_uri
      ports { container_port = 8080 }
      resources {
        limits = { cpu = "1", memory = "2Gi" }
        cpu_idle = true
        startup_cpu_boost = true
      }
      env { name = "VERTEXAI_PROJECT" value = var.project_id }
      env { name = "VERTEXAI_LOCATION" value = var.vertex_location }
      env { name = "STORE_MODEL_IN_DB" value = "True" }
      env {
        name = "LITELLM_MASTER_KEY"
        value_source { secret_key_ref { secret = google_secret_manager_secret.master_key.secret_id version = "latest" } }
      }
      env {
        name = "LITELLM_SALT_KEY"
        value_source { secret_key_ref { secret = google_secret_manager_secret.salt_key.secret_id version = "latest" } }
      }
      env {
        name = "DATABASE_URL"
        value_source { secret_key_ref { secret = google_secret_manager_secret.database_url.secret_id version = "latest" } }
      }
      env {
        name = "TYPESAFE_API_KEY"
        value_source { secret_key_ref { secret = google_secret_manager_secret.typesafe_api_key.secret_id version = "latest" } }
      }
      volume_mounts { name = "cloudsql" mount_path = "/cloudsql" }
    }
    volumes {
      name = "cloudsql"
      cloud_sql_instance { instances = [google_sql_database_instance.postgres.connection_name] }
    }
  }
  traffic {
    type = "TRAFFIC_TARGET_ALLOCATION_TYPE_LATEST"
    percent = 100
  }
  depends_on = [
    google_project_iam_member.runtime_roles,
    google_secret_manager_secret_iam_member.runtime_secret_access,
    google_secret_manager_secret_version.master_key,
    google_secret_manager_secret_version.salt_key,
    google_secret_manager_secret_version.database_url,
    google_sql_user.litellm,
  ]
}

resource "google_cloud_run_v2_service_iam_member" "public" {
  project = var.project_id
  location = google_cloud_run_v2_service.gateway.location
  name = google_cloud_run_v2_service.gateway.name
  role = "roles/run.invoker"
  member = "allUsers"
}

What infrastructure does this create?

The configuration enables the required APIs and creates the image repository. It also creates the gateway identity with its runtime permissions.

A Cloud SQL PostgreSQL instance preserves Admin UI data. Four Secret Manager secrets provide credentials to the single-instance Cloud Run service.

output "service_url" {
  description = "LiteLLM gateway base URL."
  value = google_cloud_run_v2_service.gateway.uri
}
output "admin_url" {
  description = "LiteLLM Admin UI URL."
  value = "${google_cloud_run_v2_service.gateway.uri}/ui"
}
output "image_uri" {
  description = "Artifact Registry image URI used by Cloud Run."
  value = local.image_uri
}

Why expose these outputs?

The outputs provide the gateway address, Admin UI address, and exact container image destination after Terraform creates them.

project_id = "your-project-id"

What belongs in this file?

The example shows the required project ID setting without storing your real billed project ID in a tracked file.

import csv
import json
import os
import sys
import time
import urllib.error
import urllib.request
from collections import Counter

CASES = [
    ("simple", "SIMPLE", "Reply with only the word blue."),
    ("short_explanation", "MEDIUM", "Explain in two sentences why HTTPS is safer than HTTP."),
    ("architecture", "COMPLEX", "Design a resilient event-driven order pipeline. Include idempotency, retries, and observability in four bullets."),
    ("tradeoff", "REASONING", "Reason step by step about when a team should choose queues over synchronous APIs, considering latency, failure isolation, ordering, and operations."),
]
MODELS = {"baseline": "gemini-flash", "routed": "jev-router"}

def run_case(base_url, api_key, requested_model, case_name, expected_tier, prompt):
    payload = json.dumps({"model": requested_model, "messages": [{"role": "user", "content": prompt}]}).encode("utf-8")
    request = urllib.request.Request(
        f"{base_url}/v1/chat/completions",
        data=payload,
        headers={"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"},
        method="POST",
    )
    started = time.perf_counter()
    with urllib.request.urlopen(request, timeout=120) as response:
        body = json.loads(response.read().decode("utf-8"))
        return {
            "case": case_name,
            "expected_tier": expected_tier,
            "requested_model": requested_model,
            "returned_model": body.get("model", "unknown"),
            "latency_ms": round((time.perf_counter() - started) * 1000, 1),
            "classifier_cost_header": response.headers.get("x-litellm-classifier-cost", ""),
        }

def main():
    if len(sys.argv) != 2 or sys.argv[1] not in MODELS:
        raise SystemExit("Usage: python3 benchmark.py baseline|routed")
    base_url = os.environ["LITELLM_BASE_URL"].rstrip("/")
    api_key = os.environ["LITELLM_API_KEY"]
    mode = sys.argv[1]
    rows = []
    try:
        for case_name, expected_tier, prompt in CASES:
            row = run_case(base_url, api_key, MODELS[mode], case_name, expected_tier, prompt)
            rows.append(row)
            print(f"{case_name:18} expected={expected_tier:9} returned={row['returned_model']:28} latency_ms={row['latency_ms']:8} classifier_cost={row['classifier_cost_header'] or 'n/a'}")
    except urllib.error.HTTPError as error:
        detail = error.read().decode("utf-8", errors="replace")
        print(f"HTTP {error.code}: {detail}", file=sys.stderr)
        raise SystemExit(1) from error
    output_path = f"benchmark-{mode}.csv"
    with open(output_path, "w", newline="", encoding="utf-8") as output_file:
        writer = csv.DictWriter(output_file, fieldnames=rows[0].keys())
        writer.writeheader()
        writer.writerows(rows)
    counts = Counter(row["returned_model"] for row in rows)
    average_latency = round(sum(row["latency_ms"] for row in rows) / len(rows), 1)
    print(f"\nWrote {output_path}")
    print(f"Model counts: {dict(counts)}")
    print(f"Average latency: {average_latency} ms")

if __name__ == "__main__":
    main()

What does the benchmark prepare?

The script sends four request types to the gateway. It records the returned model, latency, and classifier cost header in separate baseline or routed CSV files.

  • Open `terraform.tfvars` in the Cloud Shell Editor.
  • Replace `your-project-id` with your billed project ID: your-project-id.
  • Save every file in the `jev-litellm-gateway` folder.
  • Switch back to the Cloud Shell terminal from earlier.
  • Initialize the project and check the configuration by running these commands:
terraform init
terraform fmt -check
terraform validate

What do these checks prove?

  • `terraform init` downloads Google provider `8.5.0` and Random provider `3.9.1`.
  • `terraform fmt -check` confirms that the Terraform files use canonical formatting.
  • `terraform validate` checks the configuration structure and provider arguments.

You should see successful initialization followed by successful validation. Your gateway now has a repeatable infrastructure definition.

Terraform validation failed?

Confirm that your terminal path ends in `jev-litellm-gateway`. Check each file against the double-check tab if Terraform identifies an unexpected block or argument.

If formatting fails, compare the whitespace in the named file with the provided version.

Ask for help with the exact file and line Terraform reports: help me diagnose this Terraform validation error.

Bootstrap the repository and image

Cloud Run needs a container image before Terraform can create the service. The first targeted apply creates only the Artifact Registry repository and the TypeSafe secret container.

  • Create the repository and TypeSafe secret container by running:
terraform apply -target=google_artifact_registry_repository.images -target=google_secret_manager_secret.typesafe_api_key

Why use a targeted apply?

Cloud Run references an image that does not exist yet. This targeted apply creates the destination first so Docker has somewhere to push the image.

  • Review the resources in the Terraform plan.
  • Confirm the apply when Terraform prompts you.

You should see Terraform report that the `jev-litellm` repository and `jev-litellm-typesafe-api-key` secret were created.

Targeted apply failed?

A newly enabled Google Cloud API can take a short time to become available. Retry the same apply once if the failure mentions service activation.

If the failure mentions permissions, confirm that Cloud Shell still targets the billed project from the previous step.

Get help without sharing your credential: help me troubleshoot this targeted Terraform apply.

Your TypeSafe key is about to leave the shell environment and enter Secret Manager. The commands write it to a temporary file without printing the value, upload it, remove the file, and clear the environment variable.

  • Add the TypeSafe key as a secret version by running these commands:
printf '%s' "$TYPESAFE_API_KEY" > /tmp/typesafe-key
gcloud secrets versions add jev-litellm-typesafe-api-key --data-file=/tmp/typesafe-key
rm /tmp/typesafe-key
unset TYPESAFE_API_KEY

How is the credential handled?

`printf` writes the existing environment value without displaying it. Secret Manager receives the temporary file as a new version.

The final two commands remove both temporary copies from the active Cloud Shell session.

You should see confirmation that a new secret version was added. The key value itself stays hidden.

Secret version was not added?

Confirm that the targeted apply created `jev-litellm-typesafe-api-key`. A missing environment value also prevents the temporary file from containing the verified key.

Do not paste the credential into chat or terminal history while debugging. Help me verify this secret upload safely.

  • Read the exact image destination by running:
terraform output -raw image_uri

What does this output provide?

Terraform prints the complete Artifact Registry destination calculated from your region, project ID, repository, image name, and LiteLLM tag.

  • Record the printed destination here: your exact Artifact Registry image URI.
  • Configure Docker authentication for the `us-central1` Artifact Registry host by running:
gcloud auth configure-docker us-central1-docker.pkg.dev

What does this command change?

The command registers the Google Cloud credential helper for the regional registry host. Docker can then use your authenticated Cloud Shell identity when it pushes the image.

  • Approve the credential-helper configuration if the command asks for confirmation.

The image build can take a few minutes on its first run because Docker downloads the pinned LiteLLM base image. A stream of downloaded layers means the build is progressing.

  • Build the image and push it to the recorded destination by running:
docker build -t [[IMAGE_URI="your exact Artifact Registry image URI"]] .
docker push [[IMAGE_URI="your exact Artifact Registry image URI"]]

What happens during the build?

Docker builds from `Dockerfile` and adds `config.yaml` to the image. The `1.104.0` tag keeps the deployed LiteLLM release aligned with the Terraform image URI.

The push uploads the resulting image layers to the `jev-litellm` repository.

You should see the push finish with an image digest. That repository now contains the `litellm:1.104.0` image required by Cloud Run.

Image build or push failed?

Confirm that `Dockerfile` and `config.yaml` are saved inside `jev-litellm-gateway`. Check that the recorded image URI begins with the regional Artifact Registry hostname.

If the push cannot authenticate, rerun the Docker authentication command for the same regional hostname.

Share the non-secret Docker error for targeted help: help me troubleshoot this LiteLLM image push.

Deploy and verify the gateway

The full apply now connects the image to managed compute, persistent PostgreSQL, and secret injection. The resulting gateway keeps its Admin UI data even when Cloud Run replaces an instance.

This step creates billed resources

The apply creates usage-billed Google Cloud resources. The project targets under $5 when you destroy the stack within two hours, but the virtual-key budget does not cap Google Cloud charges.

Cloud SQL provisioning commonly takes about 15 minutes. A long period of status updates during this apply is expected.

  • Start the complete infrastructure deployment by running:
terraform apply

What does the full apply create?

Terraform creates the PostgreSQL `15` database instance, the `litellm` database, and the `litellm` user. It also creates the gateway service account with access to Vertex AI, Cloud SQL, and Secret Manager.

The apply creates the remaining secrets before deploying a public Cloud Run service. The service scales from zero to one instance and mounts the Cloud SQL connection at `/cloudsql`.

  • Review the complete resource plan when Terraform displays it.
  • Confirm the apply when Terraform prompts you.
  • Wait until Terraform reports that the apply completed.

You should see outputs for `service_url`, `admin_url`, and `image_uri`. The infrastructure is ready for a direct health check.

Deployment stopped before Cloud Run?

Retry the apply once if an API was still activating. Terraform resumes from the resources already stored in state.

A failure on the public invoker binding can indicate an organization policy that blocks `allUsers`. That policy requires help from an administrator before this public lab path can continue.

Get help with the failed Terraform resource: help me diagnose this gateway deployment failure.

Before you check the service, do you expect the readiness request to reach the new gateway successfully?

  • Test the gateway readiness endpoint by running:
curl -s "$(terraform output -raw service_url)/health/readiness"

What does this check prove?

Terraform supplies the deployed service URL. The readiness path checks whether LiteLLM has started far enough to accept gateway traffic.

The request should succeed without a connection or server error. That response proves the public Cloud Run revision is ready.

Readiness request failed?

Wait briefly if the first Cloud Run instance is still starting. Run the same readiness request again after the service finishes its cold start.

If it continues to fail, inspect the failed Cloud Run revision for a missing secret version or database connection problem.

Ask for help using the redacted response and revision status: help me troubleshoot the gateway readiness check.

  • Print the Admin UI address by running:
terraform output -raw admin_url

What does this output show?

The output combines the Cloud Run service address with `/ui`. LiteLLM serves its database-backed administration interface at that path.

  • Record the printed address here: your LiteLLM Admin UI URL.
  • Open your LiteLLM Admin UI URL in a new browser tab.

You should see the LiteLLM sign-in page at an address ending in `/ui`. That page is the first visible proof that the gateway and its persistent Admin UI are online.

Your cloud-hosted gateway is ready and its Admin UI can store configuration in PostgreSQL. Next, you will create a budgeted virtual key and measure the fixed-model baseline.

Measure the Fixed-Model Baseline

Your LiteLLM gateway is live. A successful response still leaves one question: does every prompt need the same model?

A fixed-model baseline reveals the cost and right-sizing limitation before any routing logic is added. You will measure four prompts through the gemini-flash alias.

In this step, get ready to:
  • Authenticate to the Admin UI with the stored master key.
  • Create a $1 virtual key restricted to gemini-flash.
  • Produce baseline results for all four benchmark cases.
Sign in to the Admin UI

The jev-litellm-master-key secret is the bootstrap password for the Admin UI. Secret Manager keeps the value outside your source files.

This command reveals a live credential in your terminal. The value only needs to stay visible long enough for you to copy it.

  • Switch back to the Cloud Shell terminal from the previous step.
  • Keep screenshots away from the terminal while the master key is visible.
  • Read the current master key by running this command:
gcloud secrets versions access latest --secret=jev-litellm-master-key

What does this command retrieve?

The command reads the latest stored version of jev-litellm-master-key. LiteLLM accepts this value as the initial Admin UI password.

  • Copy the returned master key without adding spaces.
  • Return to the Admin UI browser tab from the previous step.
  • Enter admin in the username field.
  • Paste the copied master key into the password field.

The sign-in form now has both required credentials.

  • Complete the sign-in from the login form.

You will see the authenticated Admin UI navigation. The available sections include Virtual Keys.

Master key rejected?

Confirm that the username is admin. Copy the newest secret value directly from the terminal without including surrounding whitespace.

If access still fails, help me troubleshoot the LiteLLM Admin UI login.

Create the benchmark key

A virtual key gives the benchmark its own model permissions and budget. Restricting this key to gemini-flash creates a controlled baseline.

  • Select Virtual Keys in the Admin UI sidebar.
  • Start creating a virtual key from the Virtual Keys page.

The virtual-key configuration form is now ready for the baseline settings.

  • Enter benchmark-key as the key name.
  • Set the allowed model to gemini-flash.

The key is now scoped to the fixed-model endpoint used by this benchmark.

  • Set the maximum budget to $1.
  • Finish creating the virtual key.
  • Copy the generated plaintext key before leaving the confirmation view.

That is your first cost guardrail in place: benchmark-key can call gemini-flash within its $1 gateway budget.

What does the $1 budget protect?

The budget limits traffic authorized through benchmark-key. It does not cap billing for Google Cloud resources or direct provider activity.

  • Return to the Cloud Shell terminal from earlier.
  • Set the gateway base URL from the existing Terraform output by running this command:
export LITELLM_BASE_URL="$(terraform output -raw service_url)"

What does this variable hold?

The Terraform output supplies the deployed gateway address. LITELLM_BASE_URL gives benchmark.py the base URL for its requests.

The next command uses hidden input. Your terminal stays blank while it waits for the copied key.

  • Load the copied virtual key into the current shell without echoing it by running this command:
  • Paste the copied key when the terminal waits on a blank line.
  • Press Enter to finish the hidden input.
read -s LITELLM_API_KEY; export LITELLM_API_KEY; echo

Why is the key input blank?

The read -s command suppresses characters while you paste the key. The command then exports the value as LITELLM_API_KEY for the benchmark process.

Your current Cloud Shell session now has the gateway URL and virtual key available without printing the key.

Did the hidden input finish?

Press Enter after pasting the key. The terminal prompt should return on a fresh line.

If the benchmark later reports an authentication problem, help me reload my LiteLLM virtual key safely.

Run the baseline benchmark

The existing benchmark.py script sends the four entries in CASES to the gateway. It records the requested model and returned model for each response.

The script also records latency and the classifier-cost header. It writes those measurements to a CSV file named benchmark-baseline.csv.

  • Select benchmark.py in the Cloud Shell Editor Explorer.
  • Read the four entries in CASES to compare their prompt complexity.

The cases move from a one-word reply to an architecture design. The final case asks for tradeoff reasoning.

A quiet terminal between rows is expected while each model response completes.

Before you run the benchmark, do you expect prompt complexity alone to change the returned model?

  • Run the fixed-model baseline from the jev-litellm-gateway folder by running this command:
python3 benchmark.py baseline

What does the baseline measure?

  • The baseline argument selects gemini-flash from the MODELS mapping.
  • Each request goes to /v1/chat/completions with the same virtual key.
  • Each result captures returned_model, latency_ms, and classifier_cost_header.

You will see four successful rows for simple, short_explanation, architecture, and tradeoff.

Every row returns gemini-flash despite the change in prompt complexity. This is the intended fixed-model shortfall.

  • Select benchmark-baseline.csv in the Cloud Shell Editor Explorer.
  • Confirm that the file contains four data rows.
  • Check that every requested_model value is gemini-flash.
  • Check that every returned_model value is gemini-flash.
  • Read the measured latency_ms value for each case.
  • Confirm that classifier_cost_header is empty for the fixed-model rows.

You now have a repeatable baseline showing that simple work and difficult work consume the same strong model. The CSV preserves the evidence for your later comparison.

Benchmark request failed?

Confirm that you are using the same Cloud Shell session where LITELLM_BASE_URL and LITELLM_API_KEY were exported. Confirm that benchmark-key can call gemini-flash.

If the request still fails, help me diagnose my fixed-model benchmark.

Your baseline now proves the right-sizing problem with measured results. Next, you will let Jev choose a suitable model tier for each request.

Add Jev Routing in the Admin UI

Your fixed-model baseline gave you a reliable comparison point. Every prompt reached Gemini 3.5 Flash regardless of its complexity.

Your AI gateway now needs a decision layer that matches each request to a suitable model. The LiteLLM Auto Router uses Jev to make that choice before dispatch.

In this step, get ready to:
  • Map the four complexity tiers to the two model aliases.
  • Configure Jev classification with its fallback controls.
  • Test the router with the routed benchmark.
Create the Jev Auto Router

Jev assigns each request to one of four built-in complexity tiers. Each tier points LiteLLM to a model alias that can handle that level of work.

  • Switch back to the LiteLLM Admin UI from the previous step.
  • Select Models + Endpoints in the left navigation.
  • Click Add Model.
  • Select the Auto Router tab.
  • Enter jev-router as the router name.

How do the tier pools work?

The router uses Jev to place each prompt into SIMPLE, MEDIUM, COMPLEX, or REASONING.

Each tier resolves to one of the model aliases already deployed with your gateway. This keeps the routing decision separate from the model connection details.

  • Assign gemini-flash-lite to the SIMPLE tier.
  • Assign gemini-flash-lite to the MEDIUM tier.
  • Assign gemini-flash to the COMPLEX tier.
  • Assign gemini-flash to the REASONING tier.
  • Set gemini-flash as the default model.
Configure and test the classifier

The classifier settings control how long LiteLLM waits for a Jev decision. They also define how the router responds if classification becomes unavailable.

  • Expand Detailed Configuration.
  • Select JEV Classifier.
  • Keep jev-latest as the classifier model.
  • Set the timeout to 3000 ms.
  • Keep the circuit breaker enabled.

What do these safeguards control?

The timeout limits how long the gateway waits for the classification result. The circuit breaker protects requests when repeated classifier calls fail.

The default-model fallback sends affected requests to gemini-flash. Your gateway can still dispatch work during a classifier interruption.

  • Set the circuit breaker cooldown to 30 seconds.
  • Choose the default-model fallback.
  • Set the context window size to 0.
  • Enable Return raw model name.

Test Routing reveals the tier decision before the router handles live benchmark traffic. Two prompts give you an early check across opposite ends of the complexity range.

  • Enter Reply with only the word blue. in the Test Routing prompt field.
  • Run Test Routing.

You should see a simple routing decision that selects gemini-flash-lite.

  • Replace the test prompt with Reason step by step about when a team should choose queues over synchronous APIs, considering latency, failure isolation, ordering, and operations..
  • Run Test Routing again.

You should see a reasoning routing decision that selects gemini-flash.

  • Save the router configuration.

That is the routing layer in place. LiteLLM now stores the Auto Router configuration in the gateway database.

Allow the router and run the benchmark

A virtual key can only call models on its allow list. Your existing benchmark key needs access to the new router before the same four prompts can test it.

  • Return to Virtual Keys in the Admin UI.
  • Select benchmark-key.
  • Add jev-router to the allowed models.
  • Keep gemini-flash in the allowed models.
  • Confirm the maximum budget remains $1.
  • Save the virtual key changes.

Your existing key can now reach the fixed model or the router. Its plaintext value remains loaded in the current shell session.

  • Return to the Cloud Shell session from the baseline run.
  • Use the terminal that is already inside jev-litellm-gateway.

Before you run the benchmark, ask yourself which tier mapping each of the four prompts will use.

  • Run the routed benchmark by using this command:
python3 benchmark.py routed

What does this command do?

  • The routed argument selects jev-router from MODELS.
  • Each response records the returned model plus its measured latency.
  • The script captures the classifier cost header when LiteLLM returns it.
  • The script writes all four rows to benchmark-routed.csv.

The terminal prints four successful rows. The script writes benchmark-routed.csv after the final response.

At least one SIMPLE or MEDIUM row should return gemini-3.5-flash-lite. At least one COMPLEX or REASONING row should return gemini-3.5-flash.

Each routed row records latency_ms. The classifier_cost_header field contains a value when LiteLLM supplies the x-litellm-classifier-cost header.

You have closed the fixed-model gap. Your gateway now uses the cheaper model for suitable work.

Did both prompt types use one model?

A single returned model can indicate that the classifier fell back or a tier mapping needs correction. Test both ends of the workload before rerunning the benchmark.

  • Return to Test Routing for jev-router.
  • Replay Reply with only the word blue..
  • Replay Reason step by step about when a team should choose queues over synchronous APIs, considering latency, failure isolation, ordering, and operations..
  • Confirm the decision cause is jev_classifier.
  • Correct any tier mapping that differs from the configuration above.
  • Save the router configuration.
  • Run the routed benchmark again from Cloud Shell.

Help me diagnose why my Jev router returns the same model for every benchmark prompt.

Your routed benchmark now records model choice, latency, and available classifier cost data. Next, you will compare both CSV files with the routing logs and spend records in the Admin UI.

Compare Routing, Latency, and Spend

Your routed AI gateway now sends each benchmark prompt to a suitable model. The next question is whether those choices improved right-sizing without hiding their cost.

The two CSV files give you a side-by-side record of model choice plus latency. The LiteLLM Admin UI connects each routed result to its classifier decision plus spend.

A working response only proves that the gateway answered. Operational evidence shows whether the classifier ran without erasing the savings.

In this step, get ready to:
  • Compare the baseline with the routed benchmark.
  • Trace a routed request through its classification records.
  • Verify spend attributed to the virtual key.
Compare the baseline and routed results

The baseline file is your control for this comparison. Matching each case across both files isolates the effect of routing.

  • Switch back to the Cloud Shell Editor from earlier.
  • Open benchmark-baseline.csv from the jev-litellm-gateway file tree.
  • Open benchmark-routed.csv in a second editor tab.

Before you compare the rows, which cases do you expect to move to the lower-cost model?

  • Match the rows using the case column.
  • Record every case where returned_model changed.
  • Calculate routed latency_ms minus baseline latency_ms for each case.
  • Record each routed classifier_cost_header value.

You should see gemini-flash in every baseline row. The routed file should contain both gemini-3.5-flash-lite results plus gemini-3.5-flash results.

How to Read the Comparison

A positive latency difference means the routed request took longer than its baseline request. A negative difference means the routed request completed faster.

The classifier_cost_header column stores the x-litellm-classifier-cost value when LiteLLM returns that header. Keep any empty values blank in your comparison.

Good progress. Your comparison now identifies the prompts that changed models plus the latency impact of each decision.

Verify routing and spend in the Admin UI

Logs connect each routed response to the Jev decision that selected its tier. Usage connects those records to spend attributed to benchmark-key.

Matching the completion to its classifier call is a fiddly part of this check. Their timestamps place both records in the same request window.

Before you inspect a routed request, which details would prove that the classifier selected its model?

  • Switch back to the LiteLLM Admin UI from earlier.
  • Select Logs in the navigation menu.
  • Narrow the visible records to requests attributed to benchmark-key.
  • Select one successful jev-router completion from the routed benchmark.
  • Inspect the request details for cause: jev_classifier.
  • Record the selected tier from the request details.
  • Record the returned model from the same request.
  • Locate the separate TypeSafe AI classifier call in the same time window.

You should see cause: jev_classifier with the selected tier plus returned model. A separate TypeSafe classifier call should appear near the completion record.

What the Logs Prove

The routing cause proves that Jev made the classification decision. The selected tier explains why LiteLLM dispatched the request to that model.

The separate classifier call records the work performed before the model completion. Together, the two records show the full routed request path.

Can't Match the Log Records?

  • Adjust the visible time window to include the routed benchmark run.
  • Compare the timestamps of the completion with nearby TypeSafe classifier calls.
  • Confirm that the completion is attributed to benchmark-key.

Help me match these LiteLLM log records.

Before you check Usage, do you expect the classifier charge plus model completion charge to appear as one spend record or two?

  • Select Usage in the Admin UI navigation menu.
  • Narrow the visible usage to benchmark-key.
  • Locate the TypeSafe classifier spend record.
  • Locate the model completion spend record.
  • Record the total spend attributed to benchmark-key.

You should see spend attributed to benchmark-key across separate classifier plus completion records. This confirms that the gateway tracked both parts of the routed workload.

Avoid Double-Counting Classifier Spend

Jev classification creates a TypeSafe spend record. The selected model creates a separate completion spend record.

The x-litellm-classifier-cost header reports the classifier charge on the routed response. Adding that value to a Usage total that already contains the classifier record would count the same charge twice.

Can't Find Spend for the Key?

  • Confirm that the Usage view is filtered to the time of the routed benchmark.
  • Confirm that the selected records are attributed to benchmark-key.
  • Compare the Usage timestamps with the routed entries in Logs.

Help me trace spend for benchmark-key.

That's the evidence connected. Your benchmark now shows model choice, latency, classifier cost, routing cause, and virtual-key spend.

Secret mission

Prove the Budget Guardrail

Your routed gateway already records spend. Give a throwaway key a tiny budget to prove that the gateway rejects further requests after the allowance is exhausted.

Clean Up Your Resources

Clean Up Your Resources

Your Google Cloud stack can continue to incur charges until you destroy it.

Decide whether to keep your resources running, pause your requests to come back later, or delete the stack entirely.

Cost warning

Cloud SQL compute remains billed while the instance runs. Its storage remains billable while it exists.

Cloud Run, Artifact Registry, Vertex AI, and TypeSafe AI remain usage billed. The $1 LiteLLM virtual-key budget limits traffic through that key. Google Cloud billing remains separate from this gateway limit.

Pause keeps the retained infrastructure billable. Delete is the safest choice when your demo is finished.

Resources you used:

  • The Cloud Run service jev-litellm-gateway.
  • The Cloud SQL instance jev-litellm-postgres. It contains the PostgreSQL database litellm plus the persisted router, virtual-key, log, and spend data.
  • The Artifact Registry repository jev-litellm. It contains the image litellm:1.104.0.
  • Four Secret Manager secrets named jev-litellm-master-key, jev-litellm-salt-key, jev-litellm-database-url, and jev-litellm-typesafe-api-key.
  • The jev-litellm service account.
  • The IAM bindings that grant runtime access to Vertex AI, Cloud SQL, and Secret Manager.
  • The jev-litellm-gateway folder in Cloud Shell. It contains your Terraform state, configuration files, and benchmark evidence.

Keep everything running

No action is needed. Choose this option while you are actively demonstrating the gateway or continuing to test its routing.

  • Keep the deployed jev-litellm-gateway stack available only while you are using it.
  • Monitor Billing in Google Cloud for continuing infrastructure charges.
  • Monitor Usage in the LiteLLM Admin UI for virtual-key spend.
  • Keep the jev-litellm-gateway folder in Cloud Shell to retain both benchmark CSV files.

Pause - I'll come back to this later

Stop sending requests while keeping the cloud stack and project files ready for a later session. Cloud SQL compute and storage continue to incur charges during this pause.

  • Stop sending benchmark requests through benchmark-key.
  • Leave the Cloud Run minimum instance count at 0.
  • Keep the jev-litellm-gateway folder in Cloud Shell.
  • Monitor Google Cloud Billing while the Cloud SQL instance remains available.

Delete - I don't want to use this again

Remove the deployed stack to stop its continuing infrastructure charges. You will also remove the local project folder after Terraform completes.

Deleting the stack is a big step. The saved Terraform state keeps the teardown scoped to resources managed by this project.

  • Switch back to the Cloud Shell terminal in the jev-litellm-gateway folder from earlier.
  • Destroy the Terraform-managed stack by running this command:
terraform destroy

What does this command remove?

Terraform reads terraform.tfstate to identify the resources created by this project. It presents a destruction plan before deleting anything.

  • The plan removes the Cloud Run service jev-litellm-gateway.
  • The plan removes the Cloud SQL instance jev-litellm-postgres with its persisted LiteLLM data.
  • The plan removes the Artifact Registry repository jev-litellm with its container image.
  • The plan removes all four Secret Manager secrets.
  • The plan removes the jev-litellm service account plus its Terraform-managed IAM bindings.
  • Review the destruction plan in the terminal.
  • Approve the destruction when Terraform asks for confirmation.

The teardown can take several minutes while Google Cloud removes the managed resources. A quiet terminal during this process does not mean the command has stopped.

  • Confirm the final Terraform summary reports completed destruction with no failed deletions.

That summary confirms the Cloud Run service, Cloud SQL instance, Artifact Registry repository, secrets, service account, and IAM bindings were removed.

Terraform destroy not finishing?

  • Confirm the Cloud Shell terminal is still inside the jev-litellm-gateway folder.
  • Confirm your Google Cloud account still has permission to delete the resources from the billed project.
  • Ask for help with the failing resource before repeating the destroy command.

The next command permanently removes Terraform state plus both benchmark CSV files from Cloud Shell.

  • Remove the local jev-litellm-gateway folder by running these commands:
cd ..
rm -rf jev-litellm-gateway

What do these commands remove?

The first command moves the terminal to the folder above jev-litellm-gateway. The second command permanently deletes the project folder.

This also removes terraform.tfstate, benchmark-baseline.csv, and benchmark-routed.csv from Cloud Shell.

  • Check the file list in Cloud Shell Editor.

You should no longer see the jev-litellm-gateway folder.

You have closed the billing loop by removing the cloud stack and its local Terraform state.

Nice Work!

Nice Work!

Your public LiteLLM 1.104.0 gateway is live on Google Cloud Run.

Its Admin UI now provides durable evidence for every routing decision.

You've learned how to:

  • Package the gateway with Docker. Reproduce its cloud infrastructure with Terraform. Persist the Admin UI state in Cloud SQL for PostgreSQL. Inject credentials through Google Secret Manager.
  • Reveal the fixed-model limitation by sending all four prompts to gemini-flash. Compare benchmark-baseline.csv against benchmark-routed.csv. Measure latency from the benchmark output. Capture classifier cost in the routed rows.
  • Configure TypeSafe AI Jev to route simple work to gemini-3.5-flash-lite. Route complex work to gemini-3.5-flash. Confirm cause: jev_classifier in LiteLLM Logs. Distinguish classifier spend from completion spend.
  • Secret Mission: Proved the budget guardrail by exhausting a temporary key's tiny allowance. Confirmed that LiteLLM returned HTTP 422. Deleted the throwaway key after the test.

Ready to quiz yourself?