Scale FastAPI on Google Cloud Run

Deploy a FastAPI container to Cloud Run and expose stateful scaling flaws.

Introduction

30 Second Summary

A web service can look dependable while every request lands in one place. A traffic burst reveals whether it still behaves correctly when several copies handle the work.

In this project, you will build a public Backend Scaling Lab using FastAPI in a Docker container on Google Cloud Run. You will make an in-memory counter fail under horizontal scaling before deploying a stateless service revision.

What You'll Build

Your finished lab opens as live interactive API documentation where a burst of requests shows exactly which Cloud Run instances handled the work.

By the end of this project, you'll have:

  • A public API whose interactive documentation opens over HTTPS. Its /health endpoint shows the active revision plus the instance identifier.
  • Scaling evidence from a concurrent burst that reaches multiple instances. Repeated low counter values reveal why process memory cannot represent service-wide state.
  • A stateless revision where every response carries a unique request ID. Production logs let you match each request to the instance that handled it. A three-instance ceiling keeps the scaling experiment bounded.
  • Secret Mission: Raise per-instance concurrency from 1 to 5. Compare how many Cloud Run instances handle the same 12-request burst.

Are there any prerequisites?

This project needs a Windows computer with Visual Studio Code, Python, and Docker Desktop already installed. You also need internet access plus a Google account with a valid payment method for Google Cloud billing verification.

Before We Start

Your first checkpoint connects the deployment work ahead to the production problem this lab exposes. Lock in what the scaling lab demonstrates before you start building.

Prepare Your Cloud Environment

Your scaling lab needs a verified local container engine before it can move to managed cloud instances. A working local engine gives you a dependable baseline for the later deployment.

The cloud side needs an isolated billing-enabled Google Cloud project. It also needs an authenticated deployment CLI so every later command targets the correct project.

In this step, get ready to:
  • Verify Docker Desktop plus Python from PowerShell.
  • Create a dedicated Google Cloud project with billing enabled.
  • Prepare Google Cloud CLI version 588.0.0 for source deployments.
Verify Docker Desktop and Python

Docker Desktop supplies the engine that builds and runs your container. Python runs the load generator later in the lab.

  • Press the Windows key to open Windows search.
  • Type Docker Desktop into the search box.
  • Press Enter to start Docker Desktop.
  • Wait until Docker Desktop reports that its engine is running.

Docker Desktop can need a short pause while its engine starts. A quiet window during startup does not mean the engine has failed.

  • Press the Windows key to open Windows search again.
  • Type Visual Studio Code into the search box.
  • Press Enter to start Visual Studio Code.
  • Open the terminal panel in Visual Studio Code.
  • Select PowerShell as the terminal shell.

Before you run the checks, what two signs would prove that the local environment is ready?

  • Verify the container engine plus the Python interpreter by running:
docker version
python -V

What do these checks prove?

  • The first command asks the Docker client to communicate with the running Docker server.
  • The second command confirms that Windows can find your Python interpreter.

You should see separate Docker Client and Server sections. You should also see a Python version.

Missing the Docker Server section?

Return to Docker Desktop if the command only prints client information. Wait for its engine to finish starting before you rerun the checks.

If Windows cannot find either command, confirm that you opened PowerShell inside Visual Studio Code. Help me diagnose my local tool checks

Create a dedicated Google Cloud project

Adding a payment method can feel like a big commitment. Eligible new Free Trial accounts require verification but do not charge during the trial.

Free usage allowances are not a hard spending cap. A dedicated project keeps this lab separate from resources you need to retain.

  • Open the official Google Cloud Free Program page in your browser.
  • Sign in with your Google account.
  • Follow the Free Trial signup flow if your account is eligible.
  • Enter a valid payment method when the signup flow requests verification.

Why use a dedicated project?

A Google Cloud project groups the service, build history, stored images, and permissions created by this lab. A dedicated project gives the cleanup step one precise deletion target.

Billing must be linked to an active account before the project can use Google Cloud resources.

  • Open the official Create projects guide.
  • Use the Go to Manage Resources link to enter the Google Cloud console.
  • Click Create Project.
  • Enter Backend Scaling Lab in the Project name field.
  • Select your active account in the Billing account field.
  • Select No organization in the Location field if that option is available.
  • Click Create.

Good progress. Your scaling lab now has an isolated cloud home that can be removed cleanly after the project.

  • Select the new project from the project drop-down at the top of the console.
  • Open the Navigation menu.
  • Select Billing.
  • Confirm that the Billing Overview page opens for an active account linked to the project.
Install and configure Google Cloud CLI

Google Cloud CLI turns the project you created in the browser into a command-line deployment target. This lab uses version 588.0.0 so the setup follows one known command surface.

The first check may report an older installation or a missing command. Choose the tab that matches the result you see.

  • Check the installed Google Cloud CLI version by running:
gcloud version

What does this check do?

The command prints the installed Google Cloud CLI version when Windows can find it. A missing command means the installer still needs to add the CLI to your system.

✔️ I see the required version

Your installed Google Cloud CLI version is 588.0.0. The CLI is ready for initialization.

ⓧ I see an older version

Your existing installation needs to be updated to version 588.0.0 before you continue.

  • Download the current Windows installer and launch it by running these commands:
(New-Object Net.WebClient).DownloadFile("https://dl.google.com/dl/cloudsdk/channels/rapid/GoogleCloudSDKInstaller.exe", "$env:Temp\GoogleCloudSDKInstaller.exe")
& $env:Temp\GoogleCloudSDKInstaller.exe

What do these commands do?

  • The first command downloads Google's signed Windows installer into your temporary files folder.
  • The second command launches the installer from that folder.
  • Complete the installer using its default options.
  • Close the PowerShell terminal after installation finishes.
  • Open a new PowerShell terminal in Visual Studio Code.
  • Confirm the updated version by running:
gcloud version

What should I see?

The version output should now identify Google Cloud CLI 588.0.0.

Still seeing the older version?

Close every open terminal before starting a fresh one. This lets Windows reload the command path updated by the installer.

If the older version remains, rerun the installer and complete the update. Help me resolve an outdated gcloud installation

ⓧ Command not found

Windows cannot find Google Cloud CLI yet. The official installer adds the command required by the rest of this lab.

  • Download the Windows installer and launch it by running these commands:
(New-Object Net.WebClient).DownloadFile("https://dl.google.com/dl/cloudsdk/channels/rapid/GoogleCloudSDKInstaller.exe", "$env:Temp\GoogleCloudSDKInstaller.exe")
& $env:Temp\GoogleCloudSDKInstaller.exe

What do these commands do?

  • The download command saves Google's signed installer in your Windows temporary files folder.
  • The launch command starts the installer from that folder.
  • Complete the installer using its default options.
  • Close the PowerShell terminal after installation finishes.
  • Open a new PowerShell terminal in Visual Studio Code.
  • Confirm the installation by running:
gcloud version

What should I see?

The output should identify Google Cloud CLI 588.0.0.

Still seeing command not found?

Restart Visual Studio Code so its terminal receives the command path created by the installer.

If the command remains unavailable, rerun the installer and confirm that it completes successfully. Help me fix a missing gcloud command

The CLI now needs permission to act as your Google account. Initialization opens browser authorization before returning you to PowerShell.

  • Initialize Google Cloud CLI by running:
gcloud init

What does initialization configure?

Initialization authenticates your Google account. It also sets the project used by later commands.

  • Select your Google account when the browser opens.
  • Approve the requested access.
  • Return to PowerShell after authorization completes.
  • Select the dedicated Backend Scaling Lab project if PowerShell lists available projects.

PowerShell returns to the prompt after initialization. Your dedicated project is now the active CLI target.

  • Enable the APIs required for source deployment by running:
gcloud services enable run.googleapis.com cloudbuild.googleapis.com

Why enable these APIs?

The Cloud Run API manages the deployed service. The Cloud Build API builds the Dockerfile from your source.

The command returns to the prompt after both services are enabled.

  • Print the active project ID by running:
gcloud config get-value project

Why capture the project ID?

The project ID uniquely identifies the cloud project targeted by your commands. Later setup commands use it to retrieve the project number and assign a build role.

  • Record the printed project ID as PROJECT_ID.
  • Retrieve the project's generated number by running:
gcloud projects describe [[PROJECT_ID="PROJECT_ID"]] --format="value(projectNumber)"

How is the project number different?

Google Cloud generates a numeric project number for internal resource identities. The next command uses that number to identify the default compute service account.

  • Record the printed project number as PROJECT_NUMBER.

The source build runs through a project service account. That identity needs the Cloud Run Builder role before it can build your service.

  • Grant the required IAM role by running:
gcloud projects add-iam-policy-binding [[PROJECT_ID="PROJECT_ID"]] --member=serviceAccount:[[PROJECT_NUMBER="PROJECT_NUMBER"]]-compute@developer.gserviceaccount.com --role=roles/run.builder

What does this permission enable?

The binding gives the project's default compute service account the builder role required for Cloud Run source deployments. It applies only inside your dedicated project.

Role propagation can take a couple of minutes. That short wait is expected while Google Cloud distributes the new policy.

Before the final check, what account plus project do you expect the CLI to report as active?

  • Verify the local engine plus the active cloud configuration by running:
docker version
gcloud auth list
gcloud config list

What should the final checks show?

  • Docker prints both Client and Server sections.
  • Google Cloud CLI marks your Google account as active.
  • The active configuration lists PROJECT_ID as the project.

That is the cloud foundation in place. Your local engine can run containers while your authenticated CLI targets the isolated billing-enabled project.

Seeing the wrong account or project?

Rerun initialization if the wrong Google account is active. Select the dedicated lab project when the CLI asks which project to use.

If the role command reports a permission problem immediately after the binding change, wait a couple of minutes before retrying it. Help me verify my gcloud configuration

Your Windows environment now has a working container engine plus an authenticated cloud target. Next up, you will build the counter-based API and run it inside Docker.

Run the Counter API Locally

Your Windows environment is ready. The next proof needs to come from the application itself.

A local Docker container gives your FastAPI service one controlled process. Its /work endpoint uses a temporary counter that appears dependable inside that single process.

In this step, get ready to:
  • Create the four files that define the containerized API.
  • Add a process-local counter to the work endpoint.
  • Build the Docker image and test the API locally.
Create the container project files

The project needs a dedicated folder so Visual Studio Code can keep the application code beside its container instructions. You will place the folder on your Desktop so it is easy to find later.

  • Press the Windows key to open Windows search.
  • Type Visual Studio Code into the search box.
  • Press Enter to open Visual Studio Code.
  • Select Terminal from the top menu.
  • Select New Terminal to open an integrated PowerShell terminal.
  • Create the backend-scaling-lab folder on your Desktop by running these commands:
Set-Location -Path "$([Environment]::GetFolderPath('Desktop'))"
New-Item -Path "backend-scaling-lab" -ItemType Directory
Set-Location -Path "backend-scaling-lab"
code .

What do these commands do?

  • The first command moves PowerShell to your Windows Desktop folder.
  • The second command creates the backend-scaling-lab project folder.
  • The third command makes that new folder the terminal's current location.
  • The final command opens the folder as a VS Code workspace.
  • Confirm that backend-scaling-lab appears at the top of the VS Code Explorer sidebar.

Project folder did not open?

Check that PowerShell printed the new folder before it ran the final command. If the folder already existed, open that existing Desktop folder in VS Code instead of creating another copy.

If VS Code does not recognize its command-line launcher, use File followed by Open Folder to select backend-scaling-lab from your Desktop.

Help me open my backend-scaling-lab folder in VS Code.

The first three files define the dependency pin, the image build, and the files Docker should ignore. Each one controls a separate part of the container.

  • Use the Explorer's new-file control to create requirements.txt.
  • Add the FastAPI dependency by pasting this line:
fastapi[standard-no-fastapi-cloud-cli]==0.143.0

What does this file do?

The dependency declaration pins FastAPI to version 0.143.0. The selected extra supplies the standard runtime tools without the FastAPI Cloud CLI.

  • Save requirements.txt.
  • Confirm that requirements.txt appears in the Explorer sidebar.

Dependency file missing?

Check that the file sits directly inside backend-scaling-lab. Confirm that Windows did not append another file extension to its name.

Help me check my requirements.txt file.

  • Use the Explorer's new-file control to create Dockerfile.
  • Define the image build by pasting this file:
FROM python:3.14.8-slim-trixie

WORKDIR /code

COPY ./requirements.txt /code/requirements.txt

RUN pip install --no-cache-dir --upgrade -r /code/requirements.txt

COPY ./main.py /code/

CMD ["fastapi", "run", "main.py", "--port", "8080"]

What does this file do?

  • The base image supplies Python through the official 3.14.8-slim-trixie image tag.
  • The image installs the pinned dependency from requirements.txt.
  • The image copies main.py into /code.
  • The final instruction starts FastAPI on port 8080 when the container runs.
  • Save Dockerfile.
  • Confirm that its filename has no extension in the Explorer sidebar.

Dockerfile showing as a text file?

Rename the file to exactly Dockerfile if Windows added a suffix. Capitalization matters because the build command searches for that filename.

Help me correct my Dockerfile name or contents.

  • Use the Explorer's new-file control to create .dockerignore.
  • Exclude local cache files by pasting these patterns:
__pycache__/
*.pyc
.venv/
.git/

What does this file do?

The ignore patterns keep Python caches, virtual environments, and Git metadata outside the Docker build context. A smaller context gives the image builder fewer unrelated files to process.

  • Save .dockerignore.
  • Confirm that the Explorer lists requirements.txt, Dockerfile, and .dockerignore inside backend-scaling-lab.

Hidden file not listed?

Confirm that the filename starts with one period. Remove any extra extension if the Explorer shows a longer name.

Help me create a correctly named .dockerignore file.

Add the temporary counter API

A process-local counter belongs to one running application process. The local container has only one process, so repeated requests share the same request_count value.

  • Use the Explorer's new-file control to create main.py.
  • Add the imports and process-level values by pasting this first chunk:
import os
import time
import uuid

from fastapi import FastAPI

app = FastAPI()
INSTANCE_ID = uuid.uuid4().hex[:8]
request_count = 0

What does this code do?

  • The imports provide environment lookup, request delays, and generated identifiers.
  • The app value creates the FastAPI application.
  • The INSTANCE_ID value identifies the running process with a short random string.
  • The request_count value starts the temporary counter at zero.
  • Save main.py.
  • Confirm that main.py appears beside the other three files in the Explorer sidebar.

Imports underlined in the editor?

VS Code may underline the FastAPI import because the package is installed inside the Docker image. The container build supplies that dependency independently of your local Python installation.

Help me distinguish a local editor warning from a Python syntax problem.

  • Place the cursor below request_count = 0 in main.py.
  • Add the health endpoint by pasting this chunk:


@app.get("/health")
def health():
    return {
        "status": "ok",
        "instance_id": INSTANCE_ID,
        "revision": os.getenv("K_REVISION", "local"),
    }

What does this endpoint do?

  • The /health route reports that the service is responding.
  • The response includes the current INSTANCE_ID so you can identify the process.
  • The revision falls back to local because the container is running outside Cloud Run.
  • Save main.py.
  • Confirm that the VS Code outline lists health().

Health function missing from the outline?

Check the indentation inside the returned dictionary. Confirm that the route decorator begins at the left edge of the file.

Help me fix the health function structure.

  • Place the cursor below the closing brace of health().
  • Add the delayed counter endpoint by pasting this chunk:


@app.get("/work")
def work(delay: float = 2.0):
    global request_count
    request_count += 1
    bounded_delay = min(max(delay, 0.0), 5.0)
    request_id = str(uuid.uuid4())
    print(
        f"request_id={request_id} instance_id={INSTANCE_ID} delay={bounded_delay}",
        flush=True,
    )
    time.sleep(bounded_delay)
    return {
        "request_id": request_id,
        "instance_id": INSTANCE_ID,
        "revision": os.getenv("K_REVISION", "local"),
        "delay": bounded_delay,
        "counter": request_count,
    }

What does the counter endpoint do?

  • The global request_count declaration lets the function update the process-level counter.
  • The bounded delay keeps the requested pause between zero seconds and five seconds.
  • Each request receives a new request_id for log matching.
  • The response returns the temporary counter beside the request diagnostics.
  • Save main.py.
  • Confirm that the VS Code outline now lists health() and work().

Work function showing a syntax problem?

Check that every opening parenthesis or brace has a matching closing character. Confirm that global request_count remains indented inside work().

Help me fix the work function syntax.

Use the tabs below to compare every project file before you build the image.

✔️ Awesome, I've got everything!

Great. Save every file before moving to the Docker build.

ⓧ I'd like to double check the full code

fastapi[standard-no-fastapi-cloud-cli]==0.143.0

Dependency checkpoint

This file should contain one pinned dependency declaration.

FROM python:3.14.8-slim-trixie

WORKDIR /code

COPY ./requirements.txt /code/requirements.txt

RUN pip install --no-cache-dir --upgrade -r /code/requirements.txt

COPY ./main.py /code/

CMD ["fastapi", "run", "main.py", "--port", "8080"]

Image checkpoint

This file should install the dependency before copying the application code.

__pycache__/
*.pyc
.venv/
.git/

Build context checkpoint

This file should contain four ignore patterns.

import os
import time
import uuid

from fastapi import FastAPI

app = FastAPI()
INSTANCE_ID = uuid.uuid4().hex[:8]
request_count = 0


@app.get("/health")
def health():
    return {
        "status": "ok",
        "instance_id": INSTANCE_ID,
        "revision": os.getenv("K_REVISION", "local"),
    }


@app.get("/work")
def work(delay: float = 2.0):
    global request_count
    request_count += 1
    bounded_delay = min(max(delay, 0.0), 5.0)
    request_id = str(uuid.uuid4())
    print(
        f"request_id={request_id} instance_id={INSTANCE_ID} delay={bounded_delay}",
        flush=True,
    )
    time.sleep(bounded_delay)
    return {
        "request_id": request_id,
        "instance_id": INSTANCE_ID,
        "revision": os.getenv("K_REVISION", "local"),
        "delay": bounded_delay,
        "counter": request_count,
    }

Application checkpoint

This temporary version should contain both endpoints plus the process-local counter.

Build and test the container

A Docker image packages the runtime, dependency, and application code into one reproducible artifact. Building it now confirms that the four files work together.

  • Return to the PowerShell terminal in the backend-scaling-lab VS Code workspace.
  • Build the local image by running this command:
docker build -t backend-scaling-lab .

What does this command do?

Docker reads the Dockerfile from the current folder. It tags the completed image as backend-scaling-lab so the run command can identify it.

The first build can take longer because Docker may need to download the Python image. The progress lines show that the build is still moving.

  • Wait for the build command to return to the PowerShell prompt.
  • Confirm that the final build output reports a completed image without an error.

Docker build failed?

Confirm that Docker Desktop is still running. Check that PowerShell is inside backend-scaling-lab before rerunning the build.

A file-related failure usually means one of the four filenames differs from the checkpoint. Compare the Explorer list with the full-code tab.

Help me diagnose my Docker build failure.

The run command keeps control of the terminal while the API is active. You will stop it with Ctrl+C after testing.

  • Start the temporary counter API by running this command:
docker run --rm --name backend-scaling-lab -p 8080:8080 backend-scaling-lab

What does this command do?

  • The command starts a container named backend-scaling-lab from the image you built.
  • Port 8080 on Windows forwards requests to port 8080 inside the container.
  • The container removes itself after it stops because the command includes --rm.

You should see status set to ok. The response also shows an eight-character instance_id plus the local revision.

Before you call /work twice, do you expect one container to return the same instance identifier or a different one each time?

  • Use the interactive control for /work to send the first request with the default delay.
  • Use the same control to send a second request.

Both responses should show the same instance_id. The counter value should increase between the two responses.

  • Return to the PowerShell terminal to inspect the two diagnostic log lines.

Each log line should have a different request_id. Both lines should have the same instance_id.

Local API not responding?

Confirm that the terminal still shows the running container. If the command stopped, read the last lines for a missing file or Python syntax problem.

If the port is already occupied, stop the other local process that is using port 8080 before rerunning the container.

Help me diagnose why the local FastAPI container is not responding.

  • Press Ctrl+C in PowerShell to stop the local container.

The PowerShell prompt should return after the container exits. That is the local baseline proven: one container served both endpoints through one process-local counter.

Your container image is ready. Next, you will deploy the same stateful design behind Cloud Run's managed autoscaling.

Deploy the Cloud Run Service

Your counter-based FastAPI service already works inside one local Docker container. That baseline proves the image can serve requests.

Google Cloud Run now places the same image behind managed autoscaling. This deployment creates the separate instances needed to test whether the counter behaves like service-wide state.

In this step, get ready to:
  • Deploy the source-built service with controlled scaling settings.
  • Record the public service URL for later tests.
  • Verify the cloud revision through the interactive API documentation.
Deploy the source-built service

A source deployment sends your existing Dockerfile to Cloud Build. Cloud Build creates the container image that Cloud Run runs.

Artifact Registry stores the built image before Cloud Run starts a new revision.

The first source deployment can take several minutes while the image is built. A quiet terminal during part of the build is expected.

This command creates resources under your billing account. The fixed maximum of three instances limits scaling exposure.

Free usage remains an allowance with no hard spending cap.

The command can pause for approval to enable Artifact Registry. A second prompt can ask to create its repository.

  • Return to the PowerShell terminal in Visual Studio Code from earlier.
  • Confirm the terminal prompt shows the backend-scaling-lab folder.
  • Approve the prompt to enable Artifact Registry if it appears.
  • Approve the repository creation prompt if it appears.
  • Create the public service with the fixed scaling controls by running this command:
gcloud run deploy backend-scaling-lab --source . --region us-central1 --allow-unauthenticated --max 3 --min 0 --concurrency 1

What does this deployment configure?

  • The --source . option sends the current folder to Cloud Build.
  • The --region us-central1 option deploys the service in the Iowa region.
  • The --allow-unauthenticated option makes the API publicly accessible.
  • The --max 3 option caps the service at three container instances.
  • The --min 0 option allows the service to scale toward zero when it receives no traffic.
  • The --concurrency 1 option allows one active request per container instance.
  • Wait for the deployment to finish.
  • Confirm the terminal prints an HTTPS service URL.

Deployment command failed?

  • Wait a couple of minutes before retrying if the roles/run.builder binding was granted recently.
  • Confirm Google Cloud CLI still targets the dedicated billing-enabled project if the deployment references an unexpected project.
  • Review the terminal output for a permission or API error before retrying.

Help me troubleshoot this Cloud Run source deployment..

That is the main cloud step complete. Your local container now has a public Cloud Run revision.

Record the public service URL

The service URL is the permanent entry point for this lab. You will reuse it when you generate concurrent traffic.

  • Copy the HTTPS service URL printed at the end of the deployment.
  • Record the URL here: SERVICE_URL.
  • Press the Windows key to open search.
  • Type the name of the web browser installed on your computer.
  • Select the browser result.
  • Click the browser address bar.
  • Paste SERVICE_URL.
  • Add /docs to the end of the URL.
  • Press Enter to load the documentation.

The interactive FastAPI documentation should load over HTTPS. Your API now has a public cloud entry point.

Documentation page not loading?

  • Confirm you copied the complete HTTPS URL from the successful deployment output.
  • Confirm /docs appears once at the end of the URL.
  • Wait briefly before refreshing if the revision has only just finished deploying.

Help me troubleshoot the deployed FastAPI documentation page..

Verify the deployed revision

The /health endpoint reports the container's INSTANCE_ID value. It also reads the revision name from K_REVISION.

Cloud Run supplies a revision name inside each deployed container. The response gives you direct evidence that the API is running in the cloud.

  • Expand the /health endpoint in the interactive documentation.
  • Click Try it out.

Before you send the request, do you expect the revision field to identify the local container or the deployed Cloud Run revision?

  • Click Execute to send the health request.

You should see status set to ok. You should also see an instance_id generated by the running container.

The revision field should contain a generated Cloud Run revision name. That proves the response came from the deployed container.

That check closes the loop. Your public FastAPI service is running from a Cloud Run revision under the scaling controls you set.

Your public revision is live under controlled scaling limits. Next, you will send a concurrent burst that exposes what each container's private counter means in practice.

Expose the Stateful Scaling Failure

Your counter-based FastAPI service is already live on Cloud Run. One container made request_count look like a reliable service-wide total.

Horizontal scaling changes the test. A concurrent Python burst sends work to separate container instances at about the same time. The returned instance_id values expose whether process memory can represent shared state.

In this step, get ready to:
  • Create a load generator that sends 12 requests concurrently.
  • Compare instance identifiers with counter values across the burst.
  • Match response identifiers with entries in the production logs.
Create the concurrent load generator

A thread pool starts several network requests without waiting for each previous request to finish. The overlapping workload gives the autoscaler a reason to distribute requests across available instances.

  • Return to the Explorer sidebar in Visual Studio Code.
  • Select the backend-scaling-lab folder.
  • Click New File in the Explorer toolbar.
  • Name the file load_test.py.
  • Add the request helper by pasting this code into load_test.py:
import concurrent.futures
import json
import sys
import urllib.request


def call_work(base_url):
    url = f"{base_url.rstrip('/')}/work?delay=2"
    with urllib.request.urlopen(url, timeout=15) as response:
        return json.loads(response.read())

What Does This Helper Do?

  • The imports provide concurrent execution, JSON decoding, command-line arguments, and HTTP requests.
  • The call_work() helper builds the /work?delay=2 URL for one request.
  • The 15-second timeout prevents one stalled request from waiting forever.
  • The returned JSON becomes a Python value that the runner can compare.
  • Save load_test.py.
  • Check the Explorer sidebar for load_test.py.

You'll see load_test.py listed beside main.py.

Can't Create or Save the Helper?

Confirm that load_test.py sits inside backend-scaling-lab. Check that the filename ends with .py.

Check that the lines inside call_work() remain indented.

Help me check my load-test helper

The helper handles one request. The main() function creates 12 workers. It also summarizes how many container instances handled the burst.

  • Place your cursor below the call_work() function.
  • Add the load-test runner by pasting this code:
def main():
    if len(sys.argv) != 2:
        raise SystemExit("Usage: python load_test.py <service-url>")

    base_url = sys.argv[1]
    with concurrent.futures.ThreadPoolExecutor(max_workers=12) as executor:
        results = list(executor.map(call_work, [base_url] * 12))

    for result in results:
        print(result)

    instance_ids = sorted({result["instance_id"] for result in results})
    print(f"Unique instances: {len(instance_ids)}")
    print(f"Instance IDs: {instance_ids}")


if __name__ == "__main__":
    main()

How Does the Runner Work?

  • The argument check requires one deployed service URL.
  • The ThreadPoolExecutor creates 12 workers for the same request helper.
  • The result loop prints every response so you can compare counter values.
  • The final two lines count the unique instance_id values that handled the workload.
  • Save load_test.py.
  • Check that the final line calls main().

Your load generator now contains the request helper and its runnable entry point.

Seeing Indentation or Syntax Problems?

Keep the argument check, thread pool, result loop, and summary lines inside main().

Keep the final condition at the left edge of the file. Its main() call stays indented beneath it.

Help me check the load-test runner

✔️ My File Matches

Your request helper and concurrent runner are in place.

ⓧ I'd Like to Double Check the Full Code

  • Compare your complete load_test.py file with this reference:
import concurrent.futures
import json
import sys
import urllib.request


def call_work(base_url):
    url = f"{base_url.rstrip('/')}/work?delay=2"
    with urllib.request.urlopen(url, timeout=15) as response:
        return json.loads(response.read())


def main():
    if len(sys.argv) != 2:
        raise SystemExit("Usage: python load_test.py <service-url>")

    base_url = sys.argv[1]
    with concurrent.futures.ThreadPoolExecutor(max_workers=12) as executor:
        results = list(executor.map(call_work, [base_url] * 12))

    for result in results:
        print(result)

    instance_ids = sorted({result["instance_id"] for result in results})
    print(f"Unique instances: {len(instance_ids)}")
    print(f"Instance IDs: {instance_ids}")


if __name__ == "__main__":
    main()
Trigger the stateful failure

Each request remains active for two seconds. With concurrency limited to one request per instance, the overlapping burst gives Cloud Run a reason to use several instances.

  • Return to the PowerShell terminal from earlier.
  • Confirm that the terminal is inside the backend-scaling-lab folder.
  • Replace SERVICE_URL in the command below with the HTTPS service URL you saved earlier.

Before you run this, do you expect one counter sequence or several isolated sequences?

  • Send the 12-request burst by running this command:
python load_test.py SERVICE_URL

What Does the Burst Prove?

You'll see more than one unique instance_id. Different instances return repeated low counter values such as 1.

Each container owns its own copy of request_count. The value cannot represent a service-wide request total.

You found the flaw the lab was designed to expose. Your service has proven that request_count belongs to one container instance.

Only Seeing One Instance?

  • Wait for the previous requests to finish.
  • Repeat the same load-test command to generate another burst.
  • Confirm that the command uses the public HTTPS service URL.

Help me investigate a single-instance load test

Match responses with service logs

The response shows what each request returned. The service log records the request_id beside the instance_id that handled it.

Before you read the logs, which two identifiers should appear together for each handled request?

  • Read the latest 30 service log entries by running this command:
gcloud run services logs read backend-scaling-lab --region us-central1 --limit 30

What Do the Logs Confirm?

Each diagnostic line contains a request_id and an instance_id. These values connect one response to the container that processed it.

Matching the identifiers confirms that several isolated processes handled the burst. Their counters remain separate even though every request reached the same public service.

  • Choose one request_id from the load-test output.
  • Find the same request_id in the service logs.
  • Compare its logged instance_id with the response value.

You'll find the same request and instance identifiers in both places. That match proves which isolated container handled each request.

Can't Find the Request IDs?

  • Repeat the load test to create fresh request entries.
  • Read the latest service logs again.
  • Compare identifiers from the newest response lines.

Help me correlate responses with Cloud Run logs

The stateful design has failed under real horizontal scaling. Next, you'll remove the shared counter assumption and give every request an independent identity.

Deploy the Stateless Revision

Your previous load test exposed the flaw: each Cloud Run instance maintained its own counter. Repeated low values proved that process memory cannot represent service-wide state.

The correction makes every request self-contained. You will remove the counter from the FastAPI service before deploying the corrected source.

In this step, get ready to:
  • Remove process-local counter state from main.py.
  • Deploy the corrected source as a new Cloud Run revision.
  • Prove each request has independent identity through the load test and logs.
Remove the process-local counter

A stateless backend keeps each request independent from mutable process memory. Any available instance can handle the request because its identity is created inside that request.

  • In the Visual Studio Code Explorer sidebar, select main.py.
  • Remove every line that contains request_count from main.py.
  • Compare the final work() endpoint with this reference:
@app.get("/work")
def work(delay: float = 2.0):
    bounded_delay = min(max(delay, 0.0), 5.0)
    request_id = str(uuid.uuid4())
    print(
        f"request_id={request_id} instance_id={INSTANCE_ID} delay={bounded_delay}",
        flush=True,
    )
    time.sleep(bounded_delay)
    return {
        "request_id": request_id,
        "instance_id": INSTANCE_ID,
        "revision": os.getenv("K_REVISION", "local"),
        "delay": bounded_delay,
    }

What Changed?

  • The endpoint still bounds the requested delay before starting the simulated work.
  • Each call generates a new request_id for request-level tracing.
  • The INSTANCE_ID remains diagnostic metadata for identifying the serving container.
  • The response contains no mutable counter state.
  • Save main.py.

✔️ Awesome, I've got everything!

Your saved main.py now uses request-scoped identity. The process-local counter is gone.

ⓧ I'd like to double check the full code

import os
import time
import uuid

from fastapi import FastAPI

app = FastAPI()
INSTANCE_ID = uuid.uuid4().hex[:8]


@app.get("/health")
def health():
    return {
        "status": "ok",
        "instance_id": INSTANCE_ID,
        "revision": os.getenv("K_REVISION", "local"),
    }


@app.get("/work")
def work(delay: float = 2.0):
    bounded_delay = min(max(delay, 0.0), 5.0)
    request_id = str(uuid.uuid4())
    print(
        f"request_id={request_id} instance_id={INSTANCE_ID} delay={bounded_delay}",
        flush=True,
    )
    time.sleep(bounded_delay)
    return {
        "request_id": request_id,
        "instance_id": INSTANCE_ID,
        "revision": os.getenv("K_REVISION", "local"),
        "delay": bounded_delay,
    }

Why This File Is Stateless

The file keeps container identity for diagnostics. Every request receives its own generated identifier without changing shared application state.

Deploy the new revision

Cloud Run source deployment asks Cloud Build to rebuild the included Dockerfile. A successful deployment creates a new revision before directing service traffic to it.

The source build may keep the terminal busy while Cloud Build creates the image. That wait means the new revision is being assembled.

  • Switch back to the PowerShell terminal in Visual Studio Code.
  • Deploy the corrected source with the existing scaling safeguards by running:
gcloud run deploy backend-scaling-lab --source . --region us-central1 --allow-unauthenticated --max 3 --min 0 --concurrency 1

What Does This Deployment Preserve?

  • The --source . option builds the source in the current backend-scaling-lab folder.
  • The --max 3 option preserves the three-instance cost safeguard.
  • The --min 0 option allows the service to scale toward zero when idle.
  • The --concurrency 1 option preserves one concurrent request per instance.
  • The --allow-unauthenticated option keeps the API publicly accessible.
  • The --region us-central1 option updates the existing regional service.

You have cleared the central flaw: the public service now runs without a process-local counter. Your existing service URL remains the entry point for the new revision.

Deployment Not Completing?

  • Wait a couple of minutes before retrying if the builder IAM binding is still propagating.
  • Confirm the PowerShell prompt is inside the backend-scaling-lab folder so --source . includes the saved Dockerfile.
  • Confirm main.py was saved before rerunning the deployment.

Help me diagnose the failed Cloud Run source deployment.

Prove the stateless revision

The same 12-request workload provides a direct comparison with the counter-based revision. Use it to evaluate whether the new revision still depends on shared process state.

Before you run the burst, consider which response fields would prove that request identity is independent.

  • Rerun the workload against your existing service by replacing SERVICE_URL with its HTTPS URL in this command:
python load_test.py SERVICE_URL

What Should You See?

  • Each of the 12 responses contains a unique request_id value.
  • The responses can contain multiple instance_id values as Cloud Run shares the workload.
  • No response contains a counter field.
  • The final summary reports how many container instances handled this particular burst.

Still Seeing Counter Values?

  • Confirm every line containing request_count was removed from main.py.
  • Confirm the deployment completed after the final file was saved.
  • Treat a single unique instance as a possible scaling result for that run. The stateless check still passes when request IDs are unique.

Help me compare my stateless load-test output with the expected fields.

Production logs connect each request identifier to the container that handled it. They provide server-side evidence for the response payloads you just inspected.

  • Read the latest service logs by running:
gcloud run services logs read backend-scaling-lab --region us-central1 --limit 30

What Do These Logs Prove?

  • Each diagnostic line includes a request_id for one request.
  • The instance_id value identifies the container that handled that request.
  • The delay value records the bounded delay used by the endpoint.
  • The log lines contain no service-wide counter state.

You have proved the fix from both sides: responses carry request-level identity while production logs trace each request to its serving instance.

Secret mission

Tune Concurrency and Compare the Result

Your stateless service can safely send each request to any available container. Raise concurrency from one to five. Rerun the identical burst to compare instance counts before restoring the original setting.

Clean Up Your Resources

Clean Up Your Resources

Your public Google Cloud Run service can receive traffic while the dedicated project exists. Decide whether to keep your resources running, pause active use to come back later, or delete them entirely.

Cost warning

Cloud Run request-based billing includes monthly free usage. Artifact Registry includes 0.5 GiB-month of free storage.

These allowances are not a hard spending cap. Source builds use Cloud Build as a separately metered resource.

The three-instance maximum limits your scaling exposure. Delete the dedicated project if you no longer need the lab.

Resources you used:

  • Dedicated Google Cloud project PROJECT_ID.
  • Cloud Run service backend-scaling-lab with its revisions in us-central1.
  • Artifact Registry repository containing the images built from your source.
  • Cloud Build history from your source deployments.
  • Project IAM binding for roles/run.builder.
  • Local backend-scaling-lab folder containing your application files.

Keep everything running

No action needed. Choose this if you want to keep testing the public scaling lab.

  • Keep the dedicated project active.
  • Leave the maximum instance setting at 3.
  • Leave the minimum instance setting at 0.
  • Leave the concurrency setting at 1.
  • Retain the local backend-scaling-lab folder.

Public requests can wake the service when it is idle. Your built images continue to use Artifact Registry storage.

Pause - I'll come back to this later

Stop sending traffic to the service while keeping the cloud resources ready for another session.

  • Stop running load_test.py against SERVICE_URL.
  • Stop opening the public service URL.
  • Leave the minimum instance setting at 0.
  • Keep the local backend-scaling-lab folder for your return.

Cloud Run compute remains request-driven while the service is idle. Artifact Registry storage remains allocated.

Delete - I don't want to use this again

Deleting the dedicated lab project is a high-impact cleanup step. It targets every cloud resource created inside that project.

The Google Cloud CLI asks for confirmation before deleting the project. The prompt gives you one last chance to verify the project ID.

  • Return to the authenticated PowerShell session from the project.
  • Replace PROJECT_ID with the dedicated lab project ID.
  • Delete the dedicated project by running this command:
gcloud projects delete PROJECT_ID

What does this command delete?

  • The command deletes the dedicated Google Cloud project.
  • Project deletion removes the backend-scaling-lab Cloud Run service with its revisions.
  • Project deletion removes the Artifact Registry repository with its stored images.
  • Project deletion removes the Cloud Build history with the project IAM binding.
  • Compare the project ID in the confirmation prompt with your dedicated lab project ID.
  • Confirm the prompt only when both project IDs match.

You should see confirmation that deletion was requested for the dedicated project. That confirmation proves the cloud cleanup has started.

Project deletion blocked?

  • Check that your authenticated Google account owns the dedicated project.
  • Confirm that PROJECT_ID matches the project configured for this lab.
  • Ask for help with the deletion error.

Cloud project deletion leaves the local backend-scaling-lab folder on your computer. Remove that folder separately if you no longer want the project files.

  • Close every file from backend-scaling-lab in VS Code.
  • Use File Explorer to locate the backend-scaling-lab folder.
  • Select the backend-scaling-lab folder.
  • Press Shift+Delete to request permanent deletion.
  • Confirm the permanent deletion prompt.

You should no longer see backend-scaling-lab in its previous location. Your local project files are now removed.

Nice Work!

Nice Work!

You did it! Your public Backend Scaling Lab turns horizontal scaling into evidence you can reproduce through responses and production logs.

You've learned how to:

  • Built a public FastAPI service inside a reproducible Docker container. Its interactive documentation lets you call /health plus /work.
  • Deployed the container through Cloud Build to a public Cloud Run revision. Configured zero minimum instances. Capped the service at three instances. Set concurrency to one request per instance.
  • Exposed a stateful scaling failure by sending 12 concurrent requests. Replaced the isolated counter with unique request IDs. Matched each request to its container instance in production logs.
  • Secret Mission: Compared the same workload at concurrency one and concurrency five. Explained the tradeoff between instance count and per-instance pressure. Restored the concurrency-one baseline.

Ready to quiz yourself?