Build an AI Bug Triage Assistant
Build a Gemini CLI that triages bug reports with JSON outputs and evals.
Introduction
30 Second Summary
Software problems rarely arrive as tidy facts. A rushed report can leave a team guessing about its urgency or where it belongs.
In this project, you will build a command-line AI Bug Triage Assistant in Python that uses the Gemini API to turn a synthetic report into structured triage data. You will also run reusable evaluation cases that expose inconsistent classifications.
What You'll Build
You'll paste a messy synthetic report into your terminal to receive clean JSON showing its summary, category, severity, and next action.
By the end of this project, you'll have:
- A command-line assistant that turns a synthetic bug report into a triage result you can inspect immediately.
- A before-and-after demonstration that reproduces the failure of free-form model text before replacing it with a predictable JSON Schema response.
- A repeatable evaluation harness that shows PASS or REVIEW beside each labeled case. Its final score summarizes the results across the full case set.
- Secret Mission: Add a deterministic gate that sends critical or unclassified reports to human review.
Are there any prerequisites?
A Google account plus basic Python familiarity are enough because the guide covers the remaining Windows setup. Use only the synthetic reports in this project because free-tier Gemini content may be used to improve Google products.
Before We Start
Before the hands-on work begins, this is your moment to commit to building an AI Bug Triage Assistant that helps you practice Gemini API integration, structured outputs, and evaluation.
Set Up the Windows Python Workspace
An unverified Python runtime can block the SDK before your assistant makes its first request. A credential saved inside code can expose access to your account.
Visual Studio Code keeps this project inside its own virtual environment. This boundary prevents the project's packages from mixing with your other Python projects.
Windows stores your Gemini API key as a user environment variable. Your source files stay free of credentials.
In this step, get ready to:
- Confirm Python 3.10 or newer with Microsoft editor support.
- Create the ai-bug-triage workspace with an active .venv environment.
- Prepare SDK access through a pinned dependency and Windows authentication.
Verify Python and editor support
The Google Gen AI SDK requires Python 3.10 or newer. The Microsoft Python extension gives VS Code the environment controls you use throughout this project.
- Press the Windows key to open Windows Search.
- Type Visual Studio Code into Windows Search.
- Select Visual Studio Code from the search results.
- Select Terminal from the top menu.
- Select New Terminal to open the VS Code terminal.
Before you check, predict whether your installed Python meets the project's minimum version.
- Check the installed Python 3 version by running this command:
py -3 --version
What does this command check?
The py launcher selects Python 3. The version output tells you whether the installed runtime can support the pinned SDK.
✔️ I see version 3.10 or higher
Python reports 3.10 or newer. Your runtime meets the project's requirement.
ⓧ I see an older version
The installed runtime cannot support the SDK version used in this project. Install a current runtime before creating the workspace environment.
- Open the official Windows guide for the Python Install Manager.
- Install Python 3.10 or newer with the Python Install Manager.
- Return to the VS Code terminal from earlier.
- Recheck the installed version by running this command:
py -3 --version
What confirms the upgrade?
The command reads the active Python 3 installation again. Continue when the output reports 3.10 or newer.
ⓧ Command not found
Windows cannot find a Python 3 runtime through the launcher. The Python Install Manager provides the supported installation path.
- Open the official Windows guide for the Python Install Manager.
- Install Python 3.10 or newer with the Python Install Manager.
- Return to the VS Code terminal from earlier.
- Check the new installation by running this command:
py -3 --version
What confirms the installation?
The command asks the Python launcher for the installed Python 3 version. A version of 3.10 or newer means the runtime is ready.
- Select the Extensions icon in the left Activity Bar.
- Search for Python in the Extensions sidebar.
- Select Python by Microsoft from the results.
- Use the installation control if the extension is absent.
- Confirm the extension page shows that Python support is enabled.
Your editor now understands Python projects and can manage their interpreters.
Python extension missing?
Check that the extension publisher is Microsoft. Reload VS Code if the extension finishes installing but its environment commands remain unavailable.
Still stuck? Help me verify Python and the Microsoft Python extension in VS Code on Windows.
Create the workspace and virtual environment
The workspace gives every project file one predictable home. The local .venv keeps this assistant's dependencies isolated inside that folder.
- Select File from the VS Code top menu.
- Select Open Folder.
- Choose Desktop in the Windows folder picker.
- Use the folder picker's new-folder control to create ai-bug-triage.
- Select the ai-bug-triage folder.
- Confirm the folder selection to open it in VS Code.
That is your project home established. The Explorer sidebar now shows ai-bug-triage as the open folder.
The environment command asks for an environment type before showing the available interpreters. It creates .venv inside the open ai-bug-triage folder.
- Press Ctrl+Shift+P to open the Command Palette.
- Select Python: Create Environment.
- Choose Venv as the environment type.
- Choose a Python 3.10 or newer interpreter.
- Wait for VS Code to create the .venv folder.
A new Windows PowerShell terminal reads the interpreter selected for this workspace. Its prompt displays (.venv) when activation succeeds.
- Press Ctrl+Shift+P to return to the Command Palette.
- Select Python: Select Interpreter.
- Choose the interpreter inside .venv.
- Close the existing terminal session in the terminal panel.
- Select Terminal from the top menu.
- Select New Terminal to start an activated PowerShell session.
The new terminal prompt begins with (.venv). This confirms that commands now install packages into the project environment.
Virtual environment not active?
Confirm that Python: Select Interpreter points to the interpreter inside .venv. Close the old terminal before creating another one.
Still missing (.venv)? Help me activate the selected .venv in the VS Code PowerShell terminal.
Install the SDK and configure authentication
A requirements file records the exact SDK release that this project expects. Anyone rebuilding the workspace can install the same dependency version.
- Use the new-file control in the Explorer sidebar to create requirements.txt inside ai-bug-triage.
- Add the pinned dependency by replacing the file contents with this line:
google-genai==2.28.0
What does this dependency pin do?
- The package name selects Google's Python SDK for the Gemini API.
- The version pin keeps every installation on release 2.28.0.
- A fixed version prevents later package changes from silently altering this project's behavior.
- Save requirements.txt in VS Code.
- Install the pinned SDK into the active environment by running this command:
pip install -r requirements.txt
What does this command install?
The command reads requirements.txt from the open workspace. The active .venv receives google-genai release 2.28.0.
The installation output names google-genai and completes without an error.
SDK installation failed?
Check that the terminal prompt starts with (.venv). Confirm that requirements.txt contains the dependency line exactly once.
Still stuck? Help me install google-genai from requirements.txt in my active VS Code environment.
✔️ Awesome, I've got everything!
Great. Double-check that requirements.txt is saved before continuing.
ⓧ I'd like to double check the full code
google-genai==2.28.0
Google AI Studio provides the API key that authenticates the SDK. The key grants access to your Gemini API account.
Keep the key out of your code
Handling a new API key can feel risky. Windows stores it outside the project.
Treat the key like a password. Never paste it into a Python file or commit it to source control.
- Open your web browser.
- Search for Google AI Studio.
- Select the official Google AI Studio result.
- Sign in with your Google account.
- Accept the terms if Google AI Studio prompts you.
- Create or view a Gemini API key.
- Copy the API key without placing it in any project file.
A Windows user environment variable makes the credential available to newly opened terminal sessions. The SDK can then read the key without receiving it from your source code.
- Press the Windows key to open Windows Search.
- Type Environment Variables into Windows Search.
- Open Environment Variables from the search results.
- Select the new-variable control under User variables.
- Enter GEMINI_API_KEY in the Variable name field.
- Paste the copied key into the Variable value field.
- Confirm the dialogs to save the user environment variable.
Why reopen the terminal?
A terminal reads user environment variables when its process starts. The existing session cannot see the new key until you replace it.
- Return to the VS Code workspace from earlier.
- Close the existing terminal session.
- Select New Terminal from the Terminal menu.
The reopened prompt starts with (.venv). It also inherits the new user environment variable.
Before you run the final check, predict whether the whole workspace is ready.
- Verify the runtime, dependency, and authentication state by running these commands:
py -3 --version
pip install -r requirements.txt
[bool]$Env:GEMINI_API_KEY
What does the final check prove?
- The first command confirms that Windows can access Python 3.10 or newer.
- The second command confirms that the pinned SDK dependency is installed in the active environment.
- The final command checks whether GEMINI_API_KEY contains a value without printing the secret.
You'll see Python 3.10 or newer. The dependency check confirms that google-genai is present.
The final line shows True without revealing the API key. That's your foundation locked in: the workspace now has isolated dependencies with terminal-based authentication.
Final check not passing?
If the final line shows False, confirm that the variable sits under User variables. Close the terminal before opening another PowerShell session.
If the dependency command fails, confirm that (.venv) appears in the prompt. Help me diagnose my final Windows Python workspace check.
Your Windows Python workspace is ready to call Gemini safely. Next, you'll send a synthetic bug report and see your first AI triage response.
Get Your First AI Triage Result
Your Windows workspace can now run Python packages inside its own virtual environment. The next goal is a live request that prints a triage result in the terminal.
The Google Gen AI SDK gives Python access to the Gemini API. You will use it to classify one synthetic checkout failure.
This first pass asks for readable headings plus bullet points. Its free-form shape has no dependable contract for downstream Python logic.
In this step, get ready to:
- Create naive_triage.py with the model configuration.
- Send the supplied synthetic checkout report to Gemini.
- Run the script to see a readable free-form triage response.
Create the first triage script
The naive_triage.py file holds the complete request path. It keeps the model choice plus synthetic test input visible in one small script.
- Use the file creation control in the VS Code sidebar to create naive_triage.py inside the open ai-bug-triage workspace.
- Paste the complete script below into naive_triage.py:
from google import genai
MODEL = "gemini-3.8-flash"
BUG_REPORT = """
Clicking Place order returns a 500 response for every test account.
Retrying and changing payment methods do not help.
""".strip()
with genai.Client() as client:
interaction = client.interactions.create(
model=MODEL,
input=(
"Triage this software bug report. Return a friendly report with "
"headings and bullet points. Do not return JSON.\n\n"
f"{BUG_REPORT}"
),
)
print(interaction.output_text)
What does this script do?
- The genai import gives the script access to the installed SDK.
- MODEL selects the stable gemini-3.8-flash endpoint.
- BUG_REPORT holds the synthetic checkout failure used for this request.
- genai.Client() creates an authenticated client from your configured environment.
- client.interactions.create() sends the prompt plus the report to the selected model.
- interaction.output_text exposes the generated response for the final print statement.
- Save naive_triage.py.
- Confirm naive_triage.py appears beside requirements.txt in the VS Code sidebar.
Cannot find the saved script?
- Confirm the file is inside the open ai-bug-triage workspace.
- Check that the filename ends with .py.
- Compare the saved script with the reference below to catch missing indentation or punctuation.
- Still stuck? Help me check why naive_triage.py is missing or incomplete in my VS Code workspace.
✔️ Awesome, I've got everything!
Your saved script now contains the complete first AI request.
ⓧ I'd like to double check the full code
from google import genai
MODEL = "gemini-3.8-flash"
BUG_REPORT = """
Clicking Place order returns a 500 response for every test account.
Retrying and changing payment methods do not help.
""".strip()
with genai.Client() as client:
interaction = client.interactions.create(
model=MODEL,
input=(
"Triage this software bug report. Return a friendly report with "
"headings and bullet points. Do not return JSON.\n\n"
f"{BUG_REPORT}"
),
)
print(interaction.output_text)
How this reference helps
This reference shows the saved file as one complete request path. Matching it prevents missing punctuation or indentation from surfacing when you run the script.
Keep the test report synthetic
Gemini API free-tier prompts plus outputs may be used to improve Google products. Synthetic reports keep customer details outside this learning exercise.
- Review the BUG_REPORT value in naive_triage.py.
- Keep the supplied checkout report unchanged for this run.
- Exclude real customer data from every free-tier test.
What crosses the API boundary?
Only the prompt string plus BUG_REPORT are included in this model request. The environment variable authenticates the client without becoming part of the prompt.
Run the synthetic triage request
The saved script is ready to make one live request. Its exact wording may vary because the model generates free-form prose.
Before you run the script, predict whether the response will use structured JSON or reader-friendly prose.
- Send the synthetic checkout report by running this command in the active terminal:
python naive_triage.py
What should you see?
You should see a readable triage response about the synthetic checkout failure. The response uses headings or bullet points instead of a structured JSON object.
Your first live AI triage result is working. The synthetic checkout failure now becomes a readable response.
No triage response?
- If authentication fails, close the current terminal.
- Create a fresh terminal in the existing workspace so Windows reloads GEMINI_API_KEY.
- If the import fails, reselect the .venv interpreter that contains google-genai==2.28.0.
- Still stuck? Help me diagnose why python naive_triage.py does not print a Gemini response.
Your first model response is live. Next, you will test whether its friendly prose can behave like application data.
Make the Text Interface Fail
Your first Gemini request now returns a readable triage report in the terminal. The next question is whether Python can use that report as structured data.
Free-form model output reaches your program as a string. A field lookup will test whether that interface supports downstream application logic.
In this step, get ready to:
- Store the model response in result.
- Attempt to read the severity field from result.
- Run the script to inspect the application boundary.
Store the response in a variable
The current print sends the response straight to the terminal. A named variable keeps the same response available for another operation.
- In naive_triage.py, scroll to the final line.
- Find this line:
print(interaction.output_text)
What does this line do?
The interaction.output_text value contains the model's free-form response. The current line sends that value directly to the terminal.
- Replace that line with the two lines below:
result = interaction.output_text
print(result)
What changed?
The result variable keeps the output available after the API call. The print(result) line preserves the readable terminal response from the previous step.
- Save naive_triage.py.
- Run the updated script from the active terminal by using this command:
python naive_triage.py
What should you see?
You should see another readable triage report with headings or bullets. This confirms that result still holds the model response.
Response not printing?
- Check that result = interaction.output_text appears above print(result).
- Confirm that the terminal prompt still shows the active .venv.
- Confirm that GEMINI_API_KEY remains available in the reopened terminal.
- Still stuck? Help me debug why result does not print the Gemini response.
Add the field lookup
A Python dictionary exposes named values through keys. The next edit tests that access pattern against result.
- In naive_triage.py, locate the final print statement.
- Confirm that the current line is:
print(result)
Why keep this line?
The existing print keeps the complete model response visible. It gives you evidence of what the model returned before the next operation runs.
- Add a new print line directly below it so the bottom of the file looks like this:
print(result)
print(result["severity"])
What does the new line test?
The first print keeps the full response visible. The second asks Python for the value stored under the severity key.
- Save naive_triage.py.
Run the boundary test
The script now performs both operations in sequence. This run gives you direct evidence about the free-form interface.
Before you run it, do you think the field lookup will succeed or fail?
- Run the boundary test from the active terminal by using this command:
python naive_triage.py
The failure is the result
You will see the readable triage report first. Python then raises a TypeError on result["severity"].
The result value is a string. It has no named severity field.
The formatting looks structured to a person. Python only receives text.
That was the intended failure. You now have proof that readable output does not provide a dependable data contract.
Not seeing the intended TypeError?
- Confirm that print(result["severity"]) appears directly below print(result).
- Save naive_triage.py before rerunning the script.
- If the request stops before printing the report, confirm that the active terminal can access GEMINI_API_KEY.
- Still stuck? Help me reproduce the intended TypeError.
- Compare naive_triage.py with the checkpoint by choosing the tab that matches your confidence:
✔️ Awesome, I've got everything!
Your saved file now preserves the free-form response before attempting the unsupported field lookup.
ⓧ I'd like to double check the full code
from google import genai
MODEL = "gemini-3.8-flash"
BUG_REPORT = """
Clicking Place order returns a 500 response for every test account.
Retrying and changing payment methods do not help.
""".strip()
with genai.Client() as client:
interaction = client.interactions.create(
model=MODEL,
input=(
"Triage this software bug report. Return a friendly report with "
"headings and bullet points. Do not return JSON.\n\n"
f"{BUG_REPORT}"
),
)
result = interaction.output_text
print(result)
print(result["severity"])
What should match?
The final three lines store the response in result. They print the full response before attempting the severity lookup.
The failure has done its job. Next, you will replace free-form prose with a predictable response that Python can parse by field.
Enforce Structured Triage Output
The last run exposed the contract gap when Python rejected a severity lookup on result. A triage workflow needs named values that code can access predictably.
A JSON Schema gives the response a defined structure. You will use Gemini to produce a Python dictionary with four required fields.
In this step, get ready to:
- Define the required response contract with JSON Schema.
- Parse the model response into a Python dictionary.
- Verify structured triage with the synthetic checkout report.
Define the JSON Schema contract
A schema limits which fields can appear in the response. Its descriptions also give the model criteria for choosing each label.
- Create triage.py in the ai-bug-triage file sidebar.
- Paste this schema scaffold into triage.py:
import json
from google import genai
MODEL = "gemini-3.8-flash"
TRIAGE_SCHEMA = {
"type": "object",
"properties": {
"summary": {
"type": "string",
"description": "A one-sentence summary of the observed software problem.",
},
"category": {
"type": "string",
"enum": ["ui", "backend", "performance", "security", "other"],
"description": "The engineering area that should receive the report.",
},
"next_action": {
"type": "string",
"description": "One concrete next investigation or reproduction action.",
},
},
"required": ["summary", "category", "severity", "next_action"],
"additionalProperties": False,
}
What does this schema scaffold do?
- The top-level object type requires a collection of named fields.
- The category enum limits classification to five engineering areas.
- The required list names every field that must appear in a result.
- The additionalProperties setting prevents unexpected fields from entering the response.
- Save triage.py.
- Check that the scaffold is valid Python by running this command in the active .venv terminal:
python triage.py
What should you see?
You will see the terminal return to the active .venv prompt without a traceback. That confirms the current schema scaffold is valid Python.
Seeing a traceback?
- Compare each opening brace in TRIAGE_SCHEMA with its closing brace.
- Check that every property block ends with a comma.
- Still stuck? Help me find the syntax issue in my TRIAGE_SCHEMA dictionary.
Severity needs stricter guidance because it controls how urgently a report should be handled. Its enum limits the possible labels while its description defines when each label applies.
- Place the cursor directly above the next_action property in TRIAGE_SCHEMA.
- Insert the severity property by pasting this code:
"severity": {
"type": "string",
"enum": ["low", "medium", "high", "critical"],
"description": (
"Use critical for security exposure, data loss, or a service-wide "
"outage; high for a blocked core workflow with no workaround; "
"medium for degraded behavior with a workaround; low for cosmetic "
"or minor behavior."
),
},
How does severity stay consistent?
- The enum restricts the result to four severity labels.
- The description connects each label to a concrete level of impact.
- The schema controls response shape. Your evaluations will measure whether the selected label is correct.
- Save triage.py.
- Check the completed schema syntax by running this command:
python triage.py
What does this check prove?
You will see the active terminal prompt return without a traceback. Python can now load the complete schema definition.
Schema syntax failing?
- Confirm the severity property sits inside the properties dictionary.
- Confirm the closing brace for severity ends with a comma.
- Still stuck? Help me place the severity property correctly in TRIAGE_SCHEMA.
Parse the model response
The triage_bug() function keeps the API call behind one reusable interface. Deterministic parsing begins when json.loads() converts the returned JSON text into a dictionary.
- Place the cursor below the closing brace of TRIAGE_SCHEMA.
- Add the API call and parsing logic by pasting this function:
def triage_bug(bug_report: str) -> dict:
with genai.Client() as client:
interaction = client.interactions.create(
model=MODEL,
input=f"Triage this software bug report:\n\n{bug_report}",
response_format={
"type": "text",
"mime_type": "application/json",
"schema": TRIAGE_SCHEMA,
},
)
output_text = interaction.output_text
if not output_text:
raise RuntimeError("Gemini returned no text output.")
return json.loads(output_text)
What does this function do?
- The bug_report parameter holds the synthetic report supplied by the application.
- The response_format configuration requests application/json.
- The schema value supplies the response contract to the model.
- The output guard stops an empty response from reaching the parser.
- The json.loads() call returns the model output as a Python dictionary.
- Save triage.py.
- Check the function syntax by running this command:
python triage.py
What should you see now?
You will see the terminal return to its prompt without a traceback. Python has loaded triage_bug() successfully.
Function syntax failing?
- Check that triage_bug() starts at the left edge of the file.
- Check that the code inside the function uses consistent four-space indentation.
- Still stuck? Help me debug the triage_bug function structure.
Run the structured assistant
The main() function collects one synthetic report from the terminal. It prints the parsed result as indented JSON.
- Place the cursor below triage_bug().
- Add the command-line workflow by pasting this code:
def main() -> None:
bug_report = input("Paste a synthetic bug report: ").strip()
if not bug_report:
print("No bug report provided.")
return
result = triage_bug(bug_report)
print("\nTriage result:")
print(json.dumps(result, indent=2))
if __name__ == "__main__":
main()
How does the command-line workflow work?
- The input() call collects a synthetic report from the terminal.
- The empty-input check ends the program when no report is supplied.
- The triage_bug() call returns the parsed dictionary.
- The json.dumps() call formats that dictionary as indented JSON.
- Save triage.py.
✔️ Awesome, I've got everything!
Your structured assistant is ready. Double-check that you saved triage.py.
ⓧ I'd like to double check the full code
import json
from google import genai
MODEL = "gemini-3.8-flash"
TRIAGE_SCHEMA = {
"type": "object",
"properties": {
"summary": {
"type": "string",
"description": "A one-sentence summary of the observed software problem.",
},
"category": {
"type": "string",
"enum": ["ui", "backend", "performance", "security", "other"],
"description": "The engineering area that should receive the report.",
},
"severity": {
"type": "string",
"enum": ["low", "medium", "high", "critical"],
"description": (
"Use critical for security exposure, data loss, or a service-wide "
"outage; high for a blocked core workflow with no workaround; "
"medium for degraded behavior with a workaround; low for cosmetic "
"or minor behavior."
),
},
"next_action": {
"type": "string",
"description": "One concrete next investigation or reproduction action.",
},
},
"required": ["summary", "category", "severity", "next_action"],
"additionalProperties": False,
}
def triage_bug(bug_report: str) -> dict:
with genai.Client() as client:
interaction = client.interactions.create(
model=MODEL,
input=f"Triage this software bug report:\n\n{bug_report}",
response_format={
"type": "text",
"mime_type": "application/json",
"schema": TRIAGE_SCHEMA,
},
)
output_text = interaction.output_text
if not output_text:
raise RuntimeError("Gemini returned no text output.")
return json.loads(output_text)
def main() -> None:
bug_report = input("Paste a synthetic bug report: ").strip()
if not bug_report:
print("No bug report provided.")
return
result = triage_bug(bug_report)
print("\nTriage result:")
print(json.dumps(result, indent=2))
if __name__ == "__main__":
main()
- Switch back to naive_triage.py.
- Select the two synthetic report lines inside BUG_REPORT.
- Copy the selected report text.
- Delete naive_triage.py from the VS Code file sidebar.
- Confirm naive_triage.py is no longer listed.
Before you run the completed assistant, what shape do you expect the model response to have?
- Start the assistant from the active .venv terminal by running this command:
python triage.py
What happens when this runs?
The script waits at the Paste a synthetic bug report: prompt. After receiving a report, it sends the request and parses the constrained response.
- Paste the copied synthetic checkout report into the terminal.
- Press Enter to submit the report.
You will see Triage result: followed by indented JSON. The object contains summary, category, severity, and next_action.
Structured result missing?
- Reopen the VS Code terminal if the current session cannot access GEMINI_API_KEY.
- Confirm .venv appears in the new terminal prompt.
- Confirm triage.py matches the full-code tab.
- Still stuck? Help me troubleshoot my structured Gemini triage response.
That response contract is now working: your assistant returns named fields that Python can parse. Next, you will test those fields across several labeled reports.
Run Repeatable Triage Evaluations
Your JSON Schema now gives every response the same fields. That solves the parsing problem from the previous step.
A single convincing classification says little about reliability. In this step, you'll build an evaluation harness that exposes semantic mismatches across several bug types.
In this step, get ready to:
- Define synthetic bug reports with expected category and severity labels.
- Compare each structured result with its expected labels.
- Run every case and inspect the aggregate score.
Define the first evaluation cases
An evaluation case pairs a fixed input with the labels you expect your application to produce. These first two cases test a security exposure and a blocked checkout workflow.
- Select the Explorer view in the Visual Studio Code Activity Bar.
- Select the New File button in the Explorer title bar.
- Enter evals.py as the file name.
- Press Enter to create the file inside ai-bug-triage.
- Build the first labeled case set by pasting this code into evals.py:
from triage import triage_bug
TEST_CASES = [
{
"name": "Cross-account billing data",
"report": (
"Changing the account_id query parameter on /profile lets a signed-in "
"user view another customer's billing address."
),
"category": "security",
"severity": "critical",
},
{
"name": "Checkout failure",
"report": (
"Clicking Place order returns a 500 response for every test account. "
"Retrying and changing payment methods do not help."
),
"category": "backend",
"severity": "high",
},
]
What Do These Cases Measure?
- The triage_bug() import reuses the schema-constrained model call from triage.py.
- Each entry in TEST_CASES stores one synthetic report with its expected category and severity.
- The two reports test whether the assistant distinguishes a critical security exposure from a high-severity backend failure.
- Save evals.py.
- Confirm that evals.py appears beside triage.py in the Explorer view.
Can't See evals.py?
- Confirm that you created the file inside the open ai-bug-triage workspace.
- Check that the file name ends with .py.
- Still stuck? Help me create evals.py beside triage.py in my Visual Studio Code workspace.
Build the comparison loop
The case labels become useful when ordinary Python compares them with the model's structured values. The loop records a match as PASS and preserves every mismatch as REVIEW.
- Add the comparison loop beneath the closing ] in evals.py by pasting this code:
def main() -> None:
passed = 0
for case in TEST_CASES:
result = triage_bug(case["report"])
matches = (
result["category"] == case["category"]
and result["severity"] == case["severity"]
)
status = "PASS" if matches else "REVIEW"
passed += int(matches)
print(f"\n{status}: {case['name']}")
print(
f" Expected: {case['category']} / {case['severity']}\n"
f" Actual: {result['category']} / {result['severity']}"
)
print(f"\nScore: {passed}/{len(TEST_CASES)}")
if __name__ == "__main__":
main()
What Does This Code Do?
- The loop sends every stored report through triage_bug().
- The matches value becomes true only when both structured labels match the expected labels.
- The passed counter turns those matches into an aggregate score.
- A REVIEW result preserves a semantic mismatch for investigation.
- Save evals.py.
Before you run the first evaluation, do you think both structured classifications will match their expected labels?
- Evaluate the first two cases by running this command in the active terminal:
python evals.py
What Should You See?
You'll see two labeled results. Each result shows PASS or REVIEW with the expected and actual labels.
The final line uses the form Score: n/2. That first run gives you a repeatable check across more than one report.
Evaluation Run Stopped?
- Confirm that evals.py and triage.py are in the same ai-bug-triage folder.
- Confirm that the active terminal prompt includes (.venv).
- Confirm that the reopened terminal can access GEMINI_API_KEY.
- Still stuck? Help me diagnose why python evals.py cannot evaluate my first two cases.
Expand coverage across four bug types
Broader coverage reveals whether the same labels behave consistently across different reports. Add a cosmetic UI issue and a degraded performance case to complete the evaluation set.
- Find the closing ] immediately above def main() -> None: in evals.py.
- Place your cursor directly above that closing ].
- Add the final two cases by pasting this code:
{
"name": "Misaligned button label",
"report": (
"On the settings page, the Save button label is two pixels lower than "
"the Cancel button, but both buttons work."
),
"category": "ui",
"severity": "low",
},
{
"name": "Slow large dashboard",
"report": (
"The dashboard takes 9 seconds to load when an account has 5,000 "
"records. It eventually loads and filtering still works."
),
"category": "performance",
"severity": "medium",
},
Why Add These Cases?
- The misaligned label tests whether a cosmetic issue remains in the low-severity UI category.
- The slow dashboard tests whether degraded behavior with a working outcome maps to medium-severity performance.
- Together, the four cases exercise distinct combinations of category and severity.
- Save evals.py.
✔️ Awesome, I've got everything!
Your four evaluation cases and comparison loop are ready. Make sure evals.py is saved.
ⓧ I'd like to double check the full code
from triage import triage_bug
TEST_CASES = [
{
"name": "Cross-account billing data",
"report": (
"Changing the account_id query parameter on /profile lets a signed-in "
"user view another customer's billing address."
),
"category": "security",
"severity": "critical",
},
{
"name": "Checkout failure",
"report": (
"Clicking Place order returns a 500 response for every test account. "
"Retrying and changing payment methods do not help."
),
"category": "backend",
"severity": "high",
},
{
"name": "Misaligned button label",
"report": (
"On the settings page, the Save button label is two pixels lower than "
"the Cancel button, but both buttons work."
),
"category": "ui",
"severity": "low",
},
{
"name": "Slow large dashboard",
"report": (
"The dashboard takes 9 seconds to load when an account has 5,000 "
"records. It eventually loads and filtering still works."
),
"category": "performance",
"severity": "medium",
},
]
def main() -> None:
passed = 0
for case in TEST_CASES:
result = triage_bug(case["report"])
matches = (
result["category"] == case["category"]
and result["severity"] == case["severity"]
)
status = "PASS" if matches else "REVIEW"
passed += int(matches)
print(f"\n{status}: {case['name']}")
print(
f" Expected: {case['category']} / {case['severity']}\n"
f" Actual: {result['category']} / {result['severity']}"
)
print(f"\nScore: {passed}/{len(TEST_CASES)}")
if __name__ == "__main__":
main()
Before you run the complete harness, which case do you think is most likely to produce a REVIEW result?
- Evaluate all four synthetic cases by running this command in the active terminal:
python evals.py
Read the Evaluation Results
You'll see PASS or REVIEW for all four cases. Each result includes the expected and actual category and severity.
The script finishes with a score in the form Score: n/4. A REVIEW result is useful evidence for improving the prompt or schema descriptions.
Missing Cases or Score?
- Confirm that all four dictionaries sit inside TEST_CASES before its closing ].
- Check that main() remains below the complete TEST_CASES list.
- Compare the indentation around for case in TEST_CASES: with the full-code tab.
- Still stuck? Help me debug missing results or an incorrect score from evals.py.
Secret mission
Add a Human Review Gate
A schema keeps each response structurally predictable, but valid fields can still carry a risky classification. In this secret mission, you will add a deterministic gate that routes critical or unclassified reports to human review.
Clean Up Your Resources
Clean Up Your Resources
The Gemini API free tier has no required ongoing cost. Decide whether to keep your resources ready, pause your work, or delete everything.
Keep Reports Synthetic
Free-tier prompts and outputs may be used to improve Google products.
- Keep real customer reports, credentials, personal data, and confidential code out of this project.
Resources you used:
- Local ai-bug-triage workspace containing .venv, requirements.txt, triage.py, and evals.py.
- Local Google Gen AI SDK installation pinned to google-genai==2.28.0 inside .venv.
- Windows user environment variable named GEMINI_API_KEY.
- Gemini API key managed through Google AI Studio.
- Associated Google Cloud project created for Gemini API access.
Keep everything running
No action is required. Choose this if you want to continue testing the assistant with synthetic reports.
- Keep the ai-bug-triage workspace for future experiments.
- Leave GEMINI_API_KEY configured while you still need API access.
- Run triage.py when you want to classify another synthetic report.
- Run evals.py when you want to measure changes against the five evaluation cases.
Pause - I'll come back to this later
Shut down running processes to prevent new API requests while keeping your files. Your API key remains active until you delete it.
- Stop any Python script that is still running in Visual Studio Code.
- Close the active terminal in Visual Studio Code.
- Close Visual Studio Code.
- Return later by reopening the existing ai-bug-triage workspace.
Delete - I don't want to use this again
Removing these resources is permanent. The project instructions let you rebuild the assistant later if you change your mind.
- Close Visual Studio Code.
- Use the Windows file browser to delete the ai-bug-triage folder.
- Open Environment Variables from Windows Search.
- Remove GEMINI_API_KEY from User variables.
- Open the Google AI Studio API Keys page.
- Delete the Gemini API key used for this project.
- Delete the associated Google Cloud project if it was created only for this exercise.
Nice Work!
Nice Work!
You did it! Your Python bug triage workflow now produces structured results through the Gemini API with evaluations and human-review safeguards.
You've learned how to:
- Integrate a foundation-model API into a Python application that classifies synthetic bug reports.
- Expose the limits of free-form model text by reproducing a failure in downstream Python logic.
- Enforce predictable model responses with JSON Schema. Build an evaluation harness that reveals semantic mismatches across labeled cases.
- Complete the optional Secret Mission by adding deterministic human-review routing for critical or unclassified reports.
Ready to quiz yourself?