Build a Guardrailed Local Python Agent
Build a local AI agent that diagnoses Python bugs and proposes safe patches.
Introduction
30 Second Summary
A coding assistant can spot a bug in seconds. Giving it permission to change files or run any command creates a much bigger risk.
In this project, you will build a local Python bug-fixing AI agent using Ollama and Qwen3 4B. The agent investigates a broken calculator before a separate approval command applies its proposed fix.
What You'll Build
Your terminal will show a local agent tracing a failed division test to the faulty line before handing you a reviewable patch.
By the end of this project, you'll have:
- Read-only investigation: Ask the agent to list and explain Python files inside sample_repo while paths outside the repository stay blocked.
- Evidence-backed diagnosis: Run the fixed unittest suite through the agent. You will see it trace 2.0 != 2.5 to the faulty division line.
- Human-controlled fixes: Preview a unified diff before a separate approval command changes calculator.py. After approval, all three tests finish with OK.
- Secret Mission: Add executable regression checks that attack path traversal, command injection, and out-of-workspace patch attempts.
Are there any prerequisites?
Your setup needs Windows 10 22H2 or newer with Python 3.8 or newer.
You also need Visual Studio Code plus internet access during setup.
The Qwen3 4B model download uses about 2.5 GB of disk space.
Before We Start
Before any hands-on work begins, take a moment to define the agent's limits. This commitment keeps source-file changes under your control throughout the project.
Set Up the Local Model and Python Environment
Your guardrails need a dependable local model server before the agent can inspect code or select tools. Ollama provides that server without sending your prompts to a hosted API.
The Qwen3 4B model supplies the local reasoning. An isolated Python virtual environment keeps the agent's dependency reproducible.
In this step, get ready to:
- Confirm Python 3.8 or newer is available.
- Verify Ollama 0.40.2 with a local Qwen3 4B response.
- Create an isolated Python environment with the Ollama SDK version 0.6.3.
Verify Python and install Ollama
The Ollama Python SDK requires Python 3.8 or newer. Start by checking which version your Windows installation exposes through the terminal.
- Press the Windows key to open the search bar.
- Type PowerShell and press Enter to open PowerShell.
- Check your installed Python version by running this command:
python --version
What does this command check?
The --version option prints the Python version linked to the python command. This confirms whether that interpreter can support the pinned Ollama SDK.
✔️ I see a supported Python version
If the output shows Python 3.8 or newer, your interpreter meets the project requirement.
ⓧ I see an older Python version
Your current interpreter is below the SDK requirement. Install a newer Python release before continuing.
- Open the official Python downloads page.
- Download a Windows installer for Python 3.8 or newer.
- Select Add Python to PATH when the installer offers that option.
- Complete the Python installation.
- Close the current PowerShell window.
- Repeat the version check in a new PowerShell window.
ⓧ The python command is unavailable
Windows cannot currently find a Python interpreter through the python command. Install Python with its command-line path enabled.
- Open the official Python downloads page.
- Download a Windows installer for Python 3.8 or newer.
- Select Add Python to PATH when the installer offers that option.
- Complete the Python installation.
- Open a new PowerShell window.
- Repeat the version check.
Ollama runs the model as a native Windows application. Its command-line tool lets you download models or start local chats from PowerShell.
- Check whether Ollama is already available by running this command:
ollama --version
What does this command check?
The version check asks the Ollama command-line tool to identify its installed release. The project targets Ollama 0.40.2.
✔️ Ollama is current
If the output names version 0.40.2, Ollama already matches the project setup.
ⓧ I see an older version
Your installed Ollama release needs an update before you download the model.
- Open the official Ollama Windows page.
- Download OllamaSetup.exe.
- Run OllamaSetup.exe from your browser downloads.
- Wait for the installer to finish.
The Windows installer updates Ollama in your user account without requiring Administrator rights.
ⓧ Command not found
Ollama is not installed yet. The official Windows installer adds its command-line tool to your user path.
- Open the official Ollama Windows page.
- Download OllamaSetup.exe.
- Run OllamaSetup.exe from your browser downloads.
- Wait for the installer to finish.
Ollama starts in the background after installation. A new PowerShell session can then find the ollama command.
Before you check the fresh installation, do you expect the new PowerShell session to recognize Ollama?
- Close the current PowerShell window.
- Press the Windows key to open the search bar.
- Type PowerShell and press Enter to open a new session.
- Verify the Ollama version and list its locally stored models by running:
ollama --version
ollama ls
What do these checks prove?
- The first command confirms that the new PowerShell session can find Ollama 0.40.2.
- The second command asks Ollama for its local model list. The list can be empty before the first model download.
You should see Ollama version 0.40.2 followed by a model table. This confirms that the local server and command-line tool are ready.
PowerShell still cannot find Ollama?
Confirm that the installer finished before opening the new PowerShell window. An older terminal keeps the path it loaded before installation.
If the new session still cannot run the command, restart Windows so the updated user path is loaded.
Help me troubleshoot Ollama after installing it on Windows.
Download and verify Qwen3 4B
The model tag qwen3:4b identifies the Qwen3 4B model in Ollama's catalog. Starting it for the first time downloads a local model artifact of about 2.5 GB.
Expect this command to take longer than the others while Ollama downloads the model. The wait depends on your internet connection and local hardware.
- Start the local model by running this command:
ollama run qwen3:4b
What does this command do?
Ollama downloads qwen3:4b when it is missing locally. It then opens an interactive chat backed by the model on your computer.
Before you send the test prompt, do you expect the reply to come from your local model without an API key?
- Send a strict response test by entering this prompt:
Reply with READY only
Why use a strict prompt?
The prompt requests one short response so the check stays easy to read. A reply proves that the downloaded model can accept local requests.
You should see a response containing READY. This is the first visible proof that local inference works.
- Press Ctrl+C to stop the interactive model session.
- Close that PowerShell window.
Your local model is responding. The agent now has a private reasoning engine ready for its tool loop.
Create the project environment
The project needs one folder for its code plus one environment for its Python package. Keeping the dependency inside tutorial-env prevents other Python projects from changing this setup.
- Press the Windows key to open the search bar.
- Type PowerShell and press Enter to open a new session.
- Move to your Desktop by running this command:
cd ~/Desktop
Why start on the Desktop?
This gives the new project folder a predictable location. You can find it later through Windows File Explorer.
- Choose this project folder name: guardrailed-local-agent.
- Create the project folder and move into it by running these commands:
mkdir [[PROJECT_FOLDER="guardrailed-local-agent"]]
cd [[PROJECT_FOLDER="guardrailed-local-agent"]]
What do these commands do?
- The first command creates your named folder on the Desktop.
- The second command makes that folder the current PowerShell location.
Your PowerShell prompt should now end with your project folder name. That confirms later setup commands will create files in the right place.
- Press the Windows key to open the search bar.
- Type Visual Studio Code and press Enter to open Visual Studio Code.
- Click File in the top menu bar.
- Click Open Folder....
- Select your guardrailed-local-agent folder from the Desktop.
- Click Select Folder.
- Confirm that you trust the folder if Visual Studio Code asks about Workspace Trust.
Why open the whole folder?
Visual Studio Code treats an opened folder as a workspace. Its Explorer can now show every file created during the project.
You should see your project folder name at the top of the Explorer sidebar. This confirms the editor is pointed at the same folder as PowerShell.
- Click the New File... button in the Explorer sidebar.
- Enter requirements.txt and press Enter.
- Add the pinned SDK dependency by pasting this file content:
ollama==0.6.3
Why pin the SDK version?
The exact ollama==0.6.3 entry makes every install use the same SDK release. This keeps the tool-calling code consistent with the project.
Using the official SDK directly keeps the upcoming agent loop focused. An extra orchestration layer would add setup without improving this project's guardrails.
- Save requirements.txt by pressing Ctrl+S.
You should see requirements.txt in the Explorer sidebar. The saved file now records the only external Python dependency.
- Switch back to the PowerShell window in your project folder.
- Create the virtual environment by running this command:
python -m venv tutorial-env
What does this command create?
Python creates an isolated interpreter plus package directory inside tutorial-env. Packages installed after activation stay inside this project environment.
- Confirm the environment folder exists by running:
ls
What should the file list show?
The directory listing should include requirements.txt plus tutorial-env. This proves the file and environment were created in the same project folder.
- Activate the virtual environment by running this command:
tutorial-env\Scripts\activate
What does activation change?
Activation points the python and pip commands at the isolated environment. Your PowerShell prompt should begin with (tutorial-env).
- Install the pinned dependency into the active environment by running:
python -m pip install -r requirements.txt
What does this install?
Pip reads requirements.txt and installs the Ollama Python SDK version 0.6.3 into tutorial-env.
You should see package installation output finish successfully. The PowerShell prompt should still begin with (tutorial-env).
Dependency installation failed?
Confirm that the prompt begins with (tutorial-env) before retrying the install. Also check that requirements.txt contains the exact pinned dependency.
Help me debug the Ollama SDK installation.
Before the final check, do you expect Python to import the SDK from tutorial-env and print the pinned version?
- Verify the installed SDK import and version by running:
python -c "import ollama; print(ollama.__version__)"
What does the final check prove?
Python imports the installed ollama package from the active environment. It then prints the package's version value.
You should see 0.6.3. That confirms the isolated Python dependency is installed at the required version.
That completes the local foundation. Ollama can serve Qwen3 4B while your activated environment can call it through the pinned Python SDK.
✔️ Awesome, I've got everything!
Great. Keep tutorial-env activated while you work through the next steps.
ⓧ I'd like to double check the full code
Compare your dependency file with this complete version. The file should contain one pinned package entry.
ollama==0.6.3
Your local model and Python environment are ready. Next, you will build the read-only agent loop that can inspect the sample repository without changing it.
Build the Read-Only Agent Loop
Your local Qwen3 model is ready. Now it needs an agent loop that can inspect real project files while staying inside a controlled workspace.
This first version exposes only read tools through Ollama. The boundary becomes visible when the agent is asked to prove the bug or apply a fix without test or write access.
In this step, get ready to:
- Create a broken calculator with three unit tests.
- Restrict file listing and reading to Python files inside sample_repo.
- Run a local agent that exposes only list_files and read_file.
Create the broken calculator fixture
A useful debugging agent needs a small repository with a real defect. The calculator uses floor division where the tests expect fractional division.
- In the VS Code Explorer sidebar, create a folder named sample_repo inside the project folder from the previous step.
- Inside sample_repo, create an empty file named __init__.py.
- Inside sample_repo, create a file named calculator.py.
- Add the deliberately broken divide() function to calculator.py by pasting this code:
def divide(total: float, groups: float) -> float:
"""Split a total evenly across a non-zero number of groups."""
if groups == 0:
raise ValueError("groups must not be zero")
return total // groups
What does this code do?
- The divide() function accepts a total and a group count.
- The zero check prevents division by zero.
- The // operator discards the fractional part of a result. This creates the bug the agent investigates.
- Save sample_repo/calculator.py.
- Check the editor diagnostics area. You should see no Python syntax problems in calculator.py.
Seeing a problem in calculator.py?
Confirm the function body uses four spaces of indentation. Check that the final line contains two forward slashes.
Help me compare my calculator function with the required code.
- Inside sample_repo, create a folder named tests.
- Inside sample_repo/tests, create an empty file named __init__.py.
- Inside sample_repo/tests, create a file named test_calculator.py.
- Add the three calculator tests to test_calculator.py by pasting this code:
import unittest
from sample_repo.calculator import divide
class DivideTests(unittest.TestCase):
def test_fractional_division(self) -> None:
self.assertEqual(divide(5, 2), 2.5)
def test_even_division(self) -> None:
self.assertEqual(divide(9, 3), 3)
def test_zero_groups(self) -> None:
with self.assertRaises(ValueError):
divide(5, 0)
if __name__ == "__main__":
unittest.main()
What do these tests cover?
- The fractional test expects divide(5, 2) to return 2.5. The current implementation returns 2.
- The even-division test confirms whole-number results still work.
- The zero-groups test confirms the function raises ValueError for an invalid divisor.
- Save sample_repo/tests/test_calculator.py.
- Confirm the Explorer sidebar lists calculator.py plus both __init__.py files plus test_calculator.py beneath sample_repo.
Missing a test file?
Expand both repository folders in the Explorer sidebar. Confirm that test_calculator.py sits inside sample_repo/tests.
Help me check my sample repository structure.
✔️ Awesome, I've got everything!
Your sample repository is ready. The calculator contains the intended division bug plus three tests that define the correct behavior.
ⓧ I'd like to double check the full code
Both sample_repo/__init__.py and sample_repo/tests/__init__.py should be empty.
def divide(total: float, groups: float) -> float:
"""Split a total evenly across a non-zero number of groups."""
if groups == 0:
raise ValueError("groups must not be zero")
return total // groups
This reference preserves the deliberate floor-division bug for the agent to investigate.
import unittest
from sample_repo.calculator import divide
class DivideTests(unittest.TestCase):
def test_fractional_division(self) -> None:
self.assertEqual(divide(5, 2), 2.5)
def test_even_division(self) -> None:
self.assertEqual(divide(9, 3), 3)
def test_zero_groups(self) -> None:
with self.assertRaises(ValueError):
divide(5, 0)
if __name__ == "__main__":
unittest.main()
This reference contains the complete test suite used by the sample repository.
Add guarded read tools
The model needs functions that translate its requests into controlled file operations. Path containment keeps every resolved path beneath sample_repo.
- In the project folder, create a file named read_tools.py.
- Add the workspace constants plus the path resolver to read_tools.py by pasting this code:
from __future__ import annotations
from pathlib import Path
PROJECT_ROOT = Path(__file__).resolve().parent
WORKSPACE_ROOT = (PROJECT_ROOT / "sample_repo").resolve()
ALLOWED_SUFFIXES = {".py"}
MAX_FILE_BYTES = 50_000
def resolve_workspace_path(relative_path: str) -> Path:
"""Resolve a path and reject anything outside the sample repository."""
candidate = (WORKSPACE_ROOT / relative_path).resolve()
try:
candidate.relative_to(WORKSPACE_ROOT)
except ValueError as exc:
raise ValueError("Path must stay inside sample_repo.") from exc
return candidate
How does path containment work?
- PROJECT_ROOT points to the folder containing the agent files.
- WORKSPACE_ROOT resolves the approved sample_repo boundary.
- resolve_workspace_path() resolves the requested path before checking whether it remains beneath that boundary.
- MAX_FILE_BYTES sets the later read limit to 50_000 bytes.
- Save read_tools.py.
- Check the editor diagnostics area. You should see no Python syntax problems in the path resolver.
Seeing a path resolver problem?
Check that Path is imported from pathlib. Confirm the exception block aligns with the try block.
Help me debug the path resolver in read_tools.py.
- Add list_files() below resolve_workspace_path() by pasting this code:
def list_files(relative_directory: str) -> dict[str, object]:
"""List Python files beneath a directory inside sample_repo.
Args:
relative_directory: Directory relative to sample_repo. Use "." for its root.
Returns:
A dictionary containing status and matching file paths.
"""
try:
directory = resolve_workspace_path(relative_directory)
if not directory.is_dir():
return {"status": "error", "error_message": "Directory does not exist."}
files = [
path.relative_to(WORKSPACE_ROOT).as_posix()
for path in sorted(directory.rglob("*.py"))
if path.is_file()
]
return {"status": "success", "files": files}
except (OSError, ValueError) as exc:
return {"status": "error", "error_message": str(exc)}
What does list_files do?
- The function passes the requested directory through the workspace guard.
- The recursive search selects only files ending in .py.
- The returned paths stay relative to sample_repo. This keeps machine-specific absolute paths out of the model conversation.
- Save read_tools.py.
- Check the editor diagnostics area. You should see no Python syntax problems around list_files().
Seeing a problem in list_files?
Check the indentation of the list comprehension. Confirm that the search pattern ends in .py.
Help me debug list_files in read_tools.py.
- Add read_file() below list_files() by pasting this code:
def read_file(relative_path: str) -> dict[str, object]:
"""Read one Python file inside sample_repo.
Args:
relative_path: Python file path relative to sample_repo.
Returns:
A dictionary containing status, path, and file content.
"""
try:
path = resolve_workspace_path(relative_path)
if path.suffix not in ALLOWED_SUFFIXES:
return {"status": "error", "error_message": "Only .py files are allowed."}
if not path.is_file():
return {"status": "error", "error_message": "File does not exist."}
if path.stat().st_size > MAX_FILE_BYTES:
return {"status": "error", "error_message": "File exceeds the read limit."}
return {
"status": "success",
"path": path.relative_to(WORKSPACE_ROOT).as_posix(),
"content": path.read_text(encoding="utf-8"),
}
except (OSError, UnicodeError, ValueError) as exc:
return {"status": "error", "error_message": str(exc)}
What does read_file protect?
- Every requested path passes through resolve_workspace_path().
- The suffix check permits only Python source files.
- The existence check rejects folders plus missing files.
- The size check keeps oversized file content out of the model context.
- Save read_tools.py.
- Check the editor diagnostics area. You should see no Python syntax problems in the completed file.
Seeing a problem in read_file?
Confirm that the returned dictionary sits inside the try block. Check that the final exception tuple contains three exception types.
Help me debug read_file in read_tools.py.
✔️ Awesome, I've got everything!
Your read layer now limits paths plus file types plus file sizes before returning content to the agent.
ⓧ I'd like to double check the full code
from __future__ import annotations
from pathlib import Path
PROJECT_ROOT = Path(__file__).resolve().parent
WORKSPACE_ROOT = (PROJECT_ROOT / "sample_repo").resolve()
ALLOWED_SUFFIXES = {".py"}
MAX_FILE_BYTES = 50_000
def resolve_workspace_path(relative_path: str) -> Path:
"""Resolve a path and reject anything outside the sample repository."""
candidate = (WORKSPACE_ROOT / relative_path).resolve()
try:
candidate.relative_to(WORKSPACE_ROOT)
except ValueError as exc:
raise ValueError("Path must stay inside sample_repo.") from exc
return candidate
def list_files(relative_directory: str) -> dict[str, object]:
"""List Python files beneath a directory inside sample_repo.
Args:
relative_directory: Directory relative to sample_repo. Use "." for its root.
Returns:
A dictionary containing status and matching file paths.
"""
try:
directory = resolve_workspace_path(relative_directory)
if not directory.is_dir():
return {"status": "error", "error_message": "Directory does not exist."}
files = [
path.relative_to(WORKSPACE_ROOT).as_posix()
for path in sorted(directory.rglob("*.py"))
if path.is_file()
]
return {"status": "success", "files": files}
except (OSError, ValueError) as exc:
return {"status": "error", "error_message": str(exc)}
def read_file(relative_path: str) -> dict[str, object]:
"""Read one Python file inside sample_repo.
Args:
relative_path: Python file path relative to sample_repo.
Returns:
A dictionary containing status, path, and file content.
"""
try:
path = resolve_workspace_path(relative_path)
if path.suffix not in ALLOWED_SUFFIXES:
return {"status": "error", "error_message": "Only .py files are allowed."}
if not path.is_file():
return {"status": "error", "error_message": "File does not exist."}
if path.stat().st_size > MAX_FILE_BYTES:
return {"status": "error", "error_message": "File exceeds the read limit."}
return {
"status": "success",
"path": path.relative_to(WORKSPACE_ROOT).as_posix(),
"content": path.read_text(encoding="utf-8"),
}
except (OSError, UnicodeError, ValueError) as exc:
return {"status": "error", "error_message": str(exc)}
This reference matches the complete guarded read layer used by the agent.
Build and test the agent loop
The loop sends the conversation plus available function tools to the model. A tool request becomes an action plus an observation before the model chooses its next step.
- In the project folder, create a file named agent.py.
- Add the imports plus the tool registry plus the system prompt to agent.py by pasting this code:
from __future__ import annotations
import json
from collections.abc import Callable
from ollama import ChatResponse, chat
from read_tools import list_files, read_file
MODEL = "qwen3:4b"
TOOL_FUNCTIONS: dict[str, Callable[..., dict[str, object]]] = {
list_files.__name__: list_files,
read_file.__name__: read_file,
}
SYSTEM_PROMPT = """
You are a careful local Python bug-fixing agent working only on sample_repo.
Paths passed to file tools are relative to sample_repo, so use "." for its root.
Use tools to inspect evidence instead of guessing. Prefer one tool call per reasoning step.
You can inspect files, but you cannot run tests or change source code.
""".strip()
How is the boundary defined?
- MODEL selects the local qwen3:4b model.
- TOOL_FUNCTIONS contains only the two guarded read functions.
- SYSTEM_PROMPT tells the model where it may inspect files. The tool registry enforces the actual capability boundary.
- Save agent.py.
- Check the editor diagnostics area. You should see no Python syntax problems in the imports or configuration.
Seeing an import problem?
Confirm the activated terminal still shows tutorial-env. Check that read_tools.py sits beside agent.py.
Help me debug the imports in agent.py.
- Add the initial run_agent_turn() structure below SYSTEM_PROMPT by pasting this code:
def run_agent_turn(messages: list[object], user_input: str) -> None:
messages.append({"role": "user", "content": user_input})
while True:
response: ChatResponse = chat(
model=MODEL,
messages=messages,
tools=list(TOOL_FUNCTIONS.values()),
)
messages.append(response.message)
print(f"Agent: {response.message.content or '(No text response)'}")
return
What does the first loop structure do?
- The user message joins the shared conversation history.
- chat() sends the model plus messages plus available tool functions to Ollama.
- The response message returns to the history before its text is printed.
- Save agent.py.
- Check the editor diagnostics area. You should see no Python syntax problems in run_agent_turn().
Seeing a problem in run_agent_turn?
Confirm that the chat() arguments align inside the call. Check that the final print remains inside the loop.
Help me debug the initial agent turn loop.
- Insert the tool-call handler between messages.append(response.message) and the final print by pasting this code:
if response.message.tool_calls:
for tool_call in response.message.tool_calls:
tool_name = tool_call.function.name
function_to_call = TOOL_FUNCTIONS.get(tool_name)
if function_to_call is None:
result = {
"status": "error",
"error_message": f"Unknown tool: {tool_name}",
}
else:
print(f"Tool: {tool_name}({dict(tool_call.function.arguments)})")
try:
result = function_to_call(**tool_call.function.arguments)
except Exception as exc:
result = {"status": "error", "error_message": str(exc)}
print(f"Observation: {json.dumps(result, ensure_ascii=False)}")
messages.append(
{
"role": "tool",
"tool_name": tool_name,
"content": json.dumps(result, ensure_ascii=False),
}
)
continue
How does the tool loop work?
- Each tool call names a function plus its arguments.
- TOOL_FUNCTIONS maps that name to an approved Python function.
- The result prints as an observation before joining the conversation as a tool message.
- continue sends the updated history back to the model until it returns ordinary text.
- Save agent.py.
- Check the editor diagnostics area. You should see no Python syntax problems in the inserted tool handler.
Seeing an indentation problem?
Place the entire handler inside the while True block. Align the new if with the final print that follows it.
Help me place the tool handler correctly.
- Add the interactive main() function below run_agent_turn() by pasting this code:
def main() -> None:
messages: list[object] = [{"role": "system", "content": SYSTEM_PROMPT}]
print("Local bug agent ready. Type quit to exit.")
while True:
user_input = input("\nYou: ").strip()
if user_input.lower() in {"quit", "exit"}:
break
if user_input:
run_agent_turn(messages, user_input)
if __name__ == "__main__":
main()
What does main do?
- The conversation starts with the system prompt.
- The input loop keeps the same message history across multiple requests.
- The values quit plus exit stop the local process.
- Save agent.py.
- Check the editor diagnostics area. You should see no Python syntax problems in the completed agent.
Seeing a problem in main?
Confirm that the final two lines start at the left edge. Check that run_agent_turn() stays inside the non-empty input condition.
Help me debug the interactive main loop.
✔️ Awesome, I've got everything!
Your agent now has a multi-turn model loop plus two approved read tools. No test runner or write function appears in its registry.
ⓧ I'd like to double check the full code
from __future__ import annotations
import json
from collections.abc import Callable
from ollama import ChatResponse, chat
from read_tools import list_files, read_file
MODEL = "qwen3:4b"
TOOL_FUNCTIONS: dict[str, Callable[..., dict[str, object]]] = {
list_files.__name__: list_files,
read_file.__name__: read_file,
}
SYSTEM_PROMPT = """
You are a careful local Python bug-fixing agent working only on sample_repo.
Paths passed to file tools are relative to sample_repo, so use "." for its root.
Use tools to inspect evidence instead of guessing. Prefer one tool call per reasoning step.
You can inspect files, but you cannot run tests or change source code.
""".strip()
def run_agent_turn(messages: list[object], user_input: str) -> None:
messages.append({"role": "user", "content": user_input})
while True:
response: ChatResponse = chat(
model=MODEL,
messages=messages,
tools=list(TOOL_FUNCTIONS.values()),
)
messages.append(response.message)
if response.message.tool_calls:
for tool_call in response.message.tool_calls:
tool_name = tool_call.function.name
function_to_call = TOOL_FUNCTIONS.get(tool_name)
if function_to_call is None:
result = {
"status": "error",
"error_message": f"Unknown tool: {tool_name}",
}
else:
print(f"Tool: {tool_name}({dict(tool_call.function.arguments)})")
try:
result = function_to_call(**tool_call.function.arguments)
except Exception as exc:
result = {"status": "error", "error_message": str(exc)}
print(f"Observation: {json.dumps(result, ensure_ascii=False)}")
messages.append(
{
"role": "tool",
"tool_name": tool_name,
"content": json.dumps(result, ensure_ascii=False),
}
)
continue
print(f"Agent: {response.message.content or '(No text response)'}")
return
def main() -> None:
messages: list[object] = [{"role": "system", "content": SYSTEM_PROMPT}]
print("Local bug agent ready. Type quit to exit.")
while True:
user_input = input("\nYou: ").strip()
if user_input.lower() in {"quit", "exit"}:
break
if user_input:
run_agent_turn(messages, user_input)
if __name__ == "__main__":
main()
This reference matches the complete read-only agent for this step.
Before you start the agent, what tool activity do you expect when it is asked to inspect the calculator?
- Start the read-only agent from the activated terminal by running this command:
python agent.py
What does this command start?
The command starts the interactive loop in agent.py using the activated environment. You should see the local-agent readiness message plus a You: prompt.
- Ask the agent to inspect the repository by entering this prompt:
List the Python files, read calculator.py, and explain divide
What should you see?
You should see tool activity for list_files plus read_file. Each action is followed by an observation containing file paths or calculator source.
You should then see an explanation of divide() based on the inspected code.
Your first local tool loop is working. The model can now gather evidence through the exact functions you approved.
Agent skipped a read tool?
Repeat the request with the required function name plus its argument. Ask for list_files with . or read_file with calculator.py.
Help me get the local agent to call its read tools.
Before the next request, do you think this tool registry can prove the defect or modify the calculator?
- Ask the read-only agent to prove the bug plus fix it by entering this prompt:
Prove the bug and fix it
Why does the agent fall short?
The agent can inspect source files. Its registry contains no function for running tests or changing code.
You should see it explain that boundary or offer an unverified suggestion. This intended shortfall proves that model-selected tools define what the loop can actually do.
- Stop the interactive loop by entering this value:
quit
What does quit do?
The input loop recognizes quit and ends the local agent process. Your project files remain unchanged.
Your read-only agent can inspect real Python code without owning test or write access. Next, you will add one fixed test command so it can replace speculation with evidence.
Add One Safe Test Command
Your read-only agent loop can inspect the calculator. However, it has no way to prove that the behavior violates a test.
This step uses command allowlisting to expose one fixed test action. The model can request that action without gaining access to a general shell.
In this step, get ready to:
- Confirm that the read-only agent cannot run the test suite.
- Create test_tool.py with one restricted run_tests function.
- Use the failing test output to diagnose the division bug.
Test the read-only boundary
The current tool registry contains only list_files plus read_file. A request for test evidence gives you a visible example of the missing capability.
Before you ask, do you think the agent can prove the failure using its current tools?
- Start the existing read-only agent from the activated terminal by running this command:
python agent.py
What does this command do?
This starts the local agent with its current two-tool registry. The process stays open so you can enter a request at the You: prompt.
- Ask the agent for test evidence by entering this prompt:
Run the sample_repo tests and diagnose the failure
Why ask for evidence first?
This request requires an actual test run. The current agent can inspect source files but has no tool that executes the tests.
You should see the agent explain that it cannot run the tests. It may inspect the source or speculate about the bug, but it cannot produce test evidence.
- Stop the current agent session by entering quit at the You: prompt.
Why this shortfall matters
The missing capability is intentional. You now have a baseline that shows the difference between reading code and proving its behavior.
Build the fixed test tool
The safe test tool stores its command as a fixed argument list. The model supplies only the permitted scope value.
- Return to Visual Studio Code from earlier.
- Click the New File control in the Explorer sidebar.
- Enter test_tool.py to create the file beside agent.py.
- Add the imports plus the fixed test argument list by copying this code into test_tool.py:
from __future__ import annotations
import subprocess
import sys
from read_tools import PROJECT_ROOT
TEST_COMMAND = [
sys.executable,
"-m",
"unittest",
"discover",
"-s",
"sample_repo/tests",
"-p",
"test_*.py",
]
What does this code do?
- The TEST_COMMAND list fixes every test-runner argument in project code.
- The sys.executable value selects the Python interpreter from your activated environment.
- The discovery arguments target the tests beneath sample_repo/tests.
- Save test_tool.py.
You should now see test_tool.py beside agent.py in the Explorer sidebar.
- Build run_tests as one function by copying the next two snippets in order.
- Place this first section beneath TEST_COMMAND:
def run_tests(test_scope: str) -> dict[str, object]:
"""Run the fixed unittest suite for the sample repository.
Args:
test_scope: Must be exactly "sample_repo".
Returns:
A dictionary containing status, return code, and captured test output.
"""
if test_scope != "sample_repo":
return {
"status": "error",
"error_message": "Only the sample_repo test scope is allowed.",
}
How does the scope gate work?
The function accepts only the literal scope sample_repo. Any other value returns an error before a subprocess starts.
- Complete run_tests by placing this second section directly beneath the scope gate:
try:
completed = subprocess.run(
TEST_COMMAND,
cwd=PROJECT_ROOT,
capture_output=True,
text=True,
timeout=30,
shell=False,
check=False,
)
except subprocess.TimeoutExpired:
return {"status": "error", "error_message": "Tests exceeded 30 seconds."}
except OSError as exc:
return {"status": "error", "error_message": str(exc)}
output = "\n".join(
part.strip() for part in (completed.stdout, completed.stderr) if part.strip()
)
return {
"status": "success" if completed.returncode == 0 else "failure",
"return_code": completed.returncode,
"command": "python -m unittest discover -s sample_repo/tests -p test_*.py",
"output": output[-12_000:],
}
What keeps the subprocess restricted?
- The subprocess receives TEST_COMMAND instead of model-generated command text.
- The cwd=PROJECT_ROOT argument runs discovery from the existing project folder.
- The shell=False argument prevents the operating system shell from interpreting the scope as a command.
- The timeout=30 argument stops a test run that exceeds the fixed limit.
- The return value captures the exit code plus the last portion of the test output.
- In agent.py, select everything from the first import through the end of SYSTEM_PROMPT.
- Replace the selected section with this updated tool registration:
from __future__ import annotations
import json
from collections.abc import Callable
from ollama import ChatResponse, chat
from read_tools import list_files, read_file
from test_tool import run_tests
MODEL = "qwen3:4b"
TOOL_FUNCTIONS: dict[str, Callable[..., dict[str, object]]] = {
list_files.__name__: list_files,
read_file.__name__: read_file,
run_tests.__name__: run_tests,
}
SYSTEM_PROMPT = """
You are a careful local Python bug-fixing agent working only on sample_repo.
Paths passed to file tools are relative to sample_repo, so use "." for its root.
Use tools to inspect evidence instead of guessing. Prefer one tool call per reasoning step.
The test tool accepts only the scope "sample_repo" and is not a general shell.
""".strip()
How does the agent gain this capability?
- The new import makes run_tests available to the agent module.
- The TOOL_FUNCTIONS entry lets the loop match the model's tool request to the Python function.
- The system prompt describes the test boundary so the model knows which scope to supply.
- No file-writing function appears in the registry.
- Save test_tool.py.
- Save agent.py.
Prove the failure with evidence
The updated agent can now ask the fixed test tool for evidence. Its result includes the test runner's exit code plus captured output.
Before you restart the agent, do you think the test suite will pass or expose the deliberate division bug?
- Start the updated agent from the activated terminal by running this command:
python agent.py
What changed in this run?
The same loop now advertises run_tests alongside its two read tools. The model still has no source-writing tool.
- Request the failing test evidence by entering this prompt:
Run the sample_repo tests and diagnose the failure
Why this prompt works
The request gives the model a test goal without giving it a shell command. The model must call run_tests with the permitted sample_repo scope.
You should see tool activity for run_tests. The captured output shows test_fractional_division failing because 2.0 != 2.5.
That failure proves divide(5, 2) uses integer-style floor division. The agent can diagnose the defect from evidence while sample_repo/calculator.py remains unchanged.
Agent skipped the test tool?
Local model behavior can vary between runs. Repeat the request with run_tests plus the required sample_repo argument stated explicitly.
If the tool returns a scope error, check that the model supplied the exact permitted value.
Help me troubleshoot the missing tool call.
✔️ Awesome, I've got everything!
Great work. Your agent can now run one fixed test suite and explain the deliberate failure without receiving general shell access.
ⓧ I'd like to double check the full code
Compare your two edited files with these exact versions.
from __future__ import annotations
import subprocess
import sys
from read_tools import PROJECT_ROOT
TEST_COMMAND = [
sys.executable,
"-m",
"unittest",
"discover",
"-s",
"sample_repo/tests",
"-p",
"test_*.py",
]
def run_tests(test_scope: str) -> dict[str, object]:
"""Run the fixed unittest suite for the sample repository.
Args:
test_scope: Must be exactly "sample_repo".
Returns:
A dictionary containing status, return code, and captured test output.
"""
if test_scope != "sample_repo":
return {
"status": "error",
"error_message": "Only the sample_repo test scope is allowed.",
}
try:
completed = subprocess.run(
TEST_COMMAND,
cwd=PROJECT_ROOT,
capture_output=True,
text=True,
timeout=30,
shell=False,
check=False,
)
except subprocess.TimeoutExpired:
return {"status": "error", "error_message": "Tests exceeded 30 seconds."}
except OSError as exc:
return {"status": "error", "error_message": str(exc)}
output = "\n".join(
part.strip() for part in (completed.stdout, completed.stderr) if part.strip()
)
return {
"status": "success" if completed.returncode == 0 else "failure",
"return_code": completed.returncode,
"command": "python -m unittest discover -s sample_repo/tests -p test_*.py",
"output": output[-12_000:],
}
What should match in test_tool.py?
Your file should contain one fixed TEST_COMMAND list. It should also contain the exact-scope gate plus the restricted subprocess call.
from __future__ import annotations
import json
from collections.abc import Callable
from ollama import ChatResponse, chat
from read_tools import list_files, read_file
from test_tool import run_tests
MODEL = "qwen3:4b"
TOOL_FUNCTIONS: dict[str, Callable[..., dict[str, object]]] = {
list_files.__name__: list_files,
read_file.__name__: read_file,
run_tests.__name__: run_tests,
}
SYSTEM_PROMPT = """
You are a careful local Python bug-fixing agent working only on sample_repo.
Paths passed to file tools are relative to sample_repo, so use "." for its root.
Use tools to inspect evidence instead of guessing. Prefer one tool call per reasoning step.
The test tool accepts only the scope "sample_repo" and is not a general shell.
""".strip()
def run_agent_turn(messages: list[object], user_input: str) -> None:
messages.append({"role": "user", "content": user_input})
while True:
response: ChatResponse = chat(
model=MODEL,
messages=messages,
tools=list(TOOL_FUNCTIONS.values()),
)
messages.append(response.message)
if response.message.tool_calls:
for tool_call in response.message.tool_calls:
tool_name = tool_call.function.name
function_to_call = TOOL_FUNCTIONS.get(tool_name)
if function_to_call is None:
result = {
"status": "error",
"error_message": f"Unknown tool: {tool_name}",
}
else:
print(f"Tool: {tool_name}({dict(tool_call.function.arguments)})")
try:
result = function_to_call(**tool_call.function.arguments)
except Exception as exc:
result = {"status": "error", "error_message": str(exc)}
print(f"Observation: {json.dumps(result, ensure_ascii=False)}")
messages.append(
{
"role": "tool",
"tool_name": tool_name,
"content": json.dumps(result, ensure_ascii=False),
}
)
continue
print(f"Agent: {response.message.content or '(No text response)'}")
return
def main() -> None:
messages: list[object] = [{"role": "system", "content": SYSTEM_PROMPT}]
print("Local bug agent ready. Type quit to exit.")
while True:
user_input = input("\nYou: ").strip()
if user_input.lower() in {"quit", "exit"}:
break
if user_input:
run_agent_turn(messages, user_input)
if __name__ == "__main__":
main()
What should match in agent.py?
Your agent should import run_tests plus register it in TOOL_FUNCTIONS. The rest of the existing loop remains unchanged.
Your agent can now prove the calculator bug through a restricted test boundary. Next, you will give it a proposal channel that records a fix without applying one.
Propose a Patch Without Applying It
Your agent can now run the fixed test suite. The failing test gives it evidence that integer division causes the defect.
A diagnosis does not grant permission to change source code. This step adds a proposal channel that records an exact replacement for later human-in-the-loop approval.
In this step, get ready to:
- Create a patch tool that validates one exact Python replacement.
- Register the proposal tool with the local agent.
- Generate a pending patch while keeping the calculator unchanged.
Create the pending patch tool
A pending patch stores the model's proposed change as data. The source file stays untouched until a separate process reviews that data.
Why use a proposal file?
The agent needs a controlled way to express intent. A proposal file captures the target path plus the exact replacement.
The original file receives a SHA-256 content hash. A later approval process can reject the proposal if the file changes after the proposal was created.
- Switch back to Visual Studio Code from earlier.
- Create a file named patch_tool.py inside the project folder from earlier.
- Add the imports plus the proposal function scaffold by pasting this code:
from __future__ import annotations
import hashlib
import json
from read_tools import PROJECT_ROOT, WORKSPACE_ROOT, resolve_workspace_path
PENDING_PATCH = PROJECT_ROOT / ".pending_patch.json"
MAX_REPLACEMENT_BYTES = 20_000
def propose_patch(
relative_path: str,
old_text: str,
new_text: str,
rationale: str,
) -> dict[str, object]:
"""Save a pending exact-text replacement without changing source code.
Args:
relative_path: Python file path relative to sample_repo.
old_text: Exact existing text to replace once.
new_text: Replacement text.
rationale: Short explanation of why the change should fix the failure.
Returns:
A dictionary describing the pending proposal or an error.
"""
What does this scaffold define?
- The imported path guard keeps patch targets beneath sample_repo.
- The PENDING_PATCH path places the proposal beside the project files. This keeps it outside the sample repository.
- The MAX_REPLACEMENT_BYTES limit prevents an oversized replacement from entering the proposal.
- The function parameters capture one target plus one exact replacement.
- Save patch_tool.py.
- Check that the scaffold loads without a traceback by running this command in the terminal from earlier:
python patch_tool.py
What does this check prove?
Python reads the new file and defines the scaffold. You should return to the PowerShell prompt without a traceback.
Seeing a syntax or import error?
Check that patch_tool.py sits beside read_tools.py. Confirm that the function docstring ends with three quotation marks.
Ask for help with the exact traceback: Help me debug why patch_tool.py cannot load beside read_tools.py.
The first guardrail checks the requested target before reading any source. It also requires the expected old text to appear exactly once.
- In propose_patch(), add this validation block immediately below the docstring:
try:
path = resolve_workspace_path(relative_path)
if path.suffix != ".py" or not path.is_file():
return {"status": "error", "error_message": "Target must be an existing .py file."}
if not old_text:
return {"status": "error", "error_message": "old_text cannot be empty."}
if len(new_text.encode("utf-8")) > MAX_REPLACEMENT_BYTES:
return {"status": "error", "error_message": "Replacement exceeds the size limit."}
current = path.read_text(encoding="utf-8")
if current.count(old_text) != 1:
return {
"status": "error",
"error_message": "old_text must match exactly once in the current file.",
}
except (OSError, UnicodeError, ValueError) as exc:
return {"status": "error", "error_message": str(exc)}
How does validation protect the workspace?
- The shared resolve_workspace_path() guard rejects targets outside sample_repo.
- The suffix check limits proposals to existing .py files.
- The exact-match check requires old_text to appear once. This avoids changing the wrong occurrence.
- The exception handler converts path failures into tool-readable error results.
A valid request now needs a durable proposal. The proposal records enough evidence for the separate approval process to verify it later.
- In propose_patch(), find the exact-match error block.
- Insert this proposal block after that error block and before the except clause:
proposal = {
"relative_path": path.relative_to(WORKSPACE_ROOT).as_posix(),
"old_text": old_text,
"new_text": new_text,
"rationale": rationale,
"sha256": hashlib.sha256(current.encode("utf-8")).hexdigest(),
}
PENDING_PATCH.write_text(
json.dumps(proposal, indent=2, ensure_ascii=False),
encoding="utf-8",
)
return {
"status": "pending",
"proposal_file": PENDING_PATCH.name,
"relative_path": proposal["relative_path"],
"message": "No source file was changed. Human approval is required.",
}
What does the proposal contain?
- The normalized relative_path identifies the approved workspace target.
- The old_text plus new_text fields describe one exact replacement.
- The rationale field records why the proposed change should fix the failure.
- The sha256 field binds the proposal to the current file contents.
- Save patch_tool.py.
- Check that the completed tool loads without a traceback by running this command:
python patch_tool.py
What should happen now?
You should return to the PowerShell prompt without a traceback. The tool is ready to receive a proposal through the agent.
Still seeing a syntax error?
Confirm that the proposal block remains inside the try block. Confirm that the except clause aligns with try.
Compare each indentation level carefully. Help me fix the indentation in propose_patch without changing its guardrails.
✔️ Awesome, I've got everything!
Your proposal tool now validates targets and stores hash-bound patch data. Make sure patch_tool.py is saved.
- Confirm that patch_tool.py sits beside agent.py.
- Confirm that the terminal check returns without a traceback.
ⓧ I'd like to double check the full code
from __future__ import annotations
import hashlib
import json
from read_tools import PROJECT_ROOT, WORKSPACE_ROOT, resolve_workspace_path
PENDING_PATCH = PROJECT_ROOT / ".pending_patch.json"
MAX_REPLACEMENT_BYTES = 20_000
def propose_patch(
relative_path: str,
old_text: str,
new_text: str,
rationale: str,
) -> dict[str, object]:
"""Save a pending exact-text replacement without changing source code.
Args:
relative_path: Python file path relative to sample_repo.
old_text: Exact existing text to replace once.
new_text: Replacement text.
rationale: Short explanation of why the change should fix the failure.
Returns:
A dictionary describing the pending proposal or an error.
"""
try:
path = resolve_workspace_path(relative_path)
if path.suffix != ".py" or not path.is_file():
return {"status": "error", "error_message": "Target must be an existing .py file."}
if not old_text:
return {"status": "error", "error_message": "old_text cannot be empty."}
if len(new_text.encode("utf-8")) > MAX_REPLACEMENT_BYTES:
return {"status": "error", "error_message": "Replacement exceeds the size limit."}
current = path.read_text(encoding="utf-8")
if current.count(old_text) != 1:
return {
"status": "error",
"error_message": "old_text must match exactly once in the current file.",
}
proposal = {
"relative_path": path.relative_to(WORKSPACE_ROOT).as_posix(),
"old_text": old_text,
"new_text": new_text,
"rationale": rationale,
"sha256": hashlib.sha256(current.encode("utf-8")).hexdigest(),
}
PENDING_PATCH.write_text(
json.dumps(proposal, indent=2, ensure_ascii=False),
encoding="utf-8",
)
return {
"status": "pending",
"proposal_file": PENDING_PATCH.name,
"relative_path": proposal["relative_path"],
"message": "No source file was changed. Human approval is required.",
}
except (OSError, UnicodeError, ValueError) as exc:
return {"status": "error", "error_message": str(exc)}
This full-file reference combines the shared path guard with exact-match validation. It also stores the pending proposal outside sample_repo.
Register the proposal tool
The function exists locally. The model can select it only after agent.py imports it and adds it to TOOL_FUNCTIONS.
- Return to agent.py in Visual Studio Code.
- Find this import group near the top of the file:
from ollama import ChatResponse, chat
from read_tools import list_files, read_file
from test_tool import run_tests
What is already imported?
The current agent imports its read tools plus the fixed test tool. It has no reference to the proposal function yet.
- Replace that import group with this updated version:
from ollama import ChatResponse, chat
from patch_tool import propose_patch
from read_tools import list_files, read_file
from test_tool import run_tests
What changed in the imports?
The new import makes propose_patch() available to the registry. Importing the function does not execute it.
- Find the current TOOL_FUNCTIONS dictionary:
TOOL_FUNCTIONS: dict[str, Callable[..., dict[str, object]]] = {
list_files.__name__: list_files,
read_file.__name__: read_file,
run_tests.__name__: run_tests,
}
What can the model select now?
The current dictionary exposes reading plus the fixed test command. No entry records a proposed replacement.
- Replace the dictionary with this updated registry:
TOOL_FUNCTIONS: dict[str, Callable[..., dict[str, object]]] = {
list_files.__name__: list_files,
read_file.__name__: read_file,
run_tests.__name__: run_tests,
propose_patch.__name__: propose_patch,
}
What does the registry control?
The registry exposes exactly four functions to the model. It contains no function that writes into sample_repo.
- Save agent.py.
- Check that the updated registry loads by running the agent:
python agent.py
What should the startup prove?
You should see Local bug agent ready. Type quit to exit. This confirms that the new import plus registry load successfully.
- Enter quit to stop this check.
Agent fails during startup?
Check that the imported name is exactly propose_patch. Check that patch_tool.py is saved beside agent.py.
Use the traceback to identify the mismatched line. Help me debug why agent.py cannot import propose_patch.
The system prompt defines the agent's working boundary. It should require evidence plus the smallest replacement while making the approval boundary explicit.
- In agent.py, replace the entire SYSTEM_PROMPT assignment with this version:
SYSTEM_PROMPT = """
You are a careful local Python bug-fixing agent working only on sample_repo.
Paths passed to file tools are relative to sample_repo, so use "." for its root.
Use tools to inspect evidence instead of guessing. Prefer one tool call per reasoning step.
The test tool accepts only the scope "sample_repo" and is not a general shell.
When fixing a bug, run tests, read the relevant file, and propose the smallest exact replacement.
You can create a pending proposal, but you cannot apply it. Never claim that source changed.
After proposing a patch, tell the user to review it with the separate approval CLI.
""".strip()
What behavior does the prompt request?
- The agent gathers test plus file evidence before proposing a fix.
- The agent requests one exact replacement through propose_patch().
- The agent states that it cannot apply the proposal.
- The agent directs the human to a separate approval process.
- Save agent.py.
Generate and inspect the pending proposal
The agent now has enough evidence to describe a fix. Its strongest available action creates .pending_patch.json outside the protected repository.
Before you make the request, do you expect sample_repo/calculator.py to change when the proposal tool runs?
- Start the updated agent by running this command:
python agent.py
What does this command start?
Python runs the local agent loop with all four registered tools. You should see the ready message followed by the You: prompt.
- Ask the agent to gather evidence and propose the minimal fix by entering this prompt:
Run the sample_repo tests. Read calculator.py. Propose the smallest exact replacement for the fractional division failure. Use propose_patch and do not claim that you applied the change.
What should the agent do?
You should see tool activity for run_tests plus read_file. You should then see propose_patch return a pending status.
The final response should explain that human approval remains required. This language matches the actual capability boundary.
Agent answers without proposing a patch?
Local model choices can vary between runs. Repeat the request with the required function plus arguments stated explicitly.
Ask it to call propose_patch for calculator.py with the exact old and new return statements. Help me guide Qwen3 to call propose_patch with the required evidence.
The proposal is now written. The remaining check proves that the recorded intent did not become a source mutation.
- Switch back to Visual Studio Code from earlier.
- Select .pending_patch.json from the project file list.
- Confirm that relative_path contains calculator.py.
- Confirm that old_text contains return total // groups.
- Confirm that new_text contains return total / groups.
- Confirm that rationale explains why true division fixes the failing test.
- Confirm that sha256 contains the original file hash.
- Select sample_repo/calculator.py from the project file list.
You should still see return total // groups in the calculator. The proposal exists while the buggy source remains unchanged.
The approval boundary is working
The model selected a tool that writes proposal data outside the protected repository. No registered tool can turn that proposal into a source change.
That is the control boundary you set out to build. The agent can recommend a fix while the final write remains yours.
✔️ Awesome, I've got everything!
Your agent can now create a hash-bound pending proposal. The calculator still uses integer division until you approve the change separately.
ⓧ I'd like to double check the full code
from __future__ import annotations
import json
from collections.abc import Callable
from ollama import ChatResponse, chat
from patch_tool import propose_patch
from read_tools import list_files, read_file
from test_tool import run_tests
MODEL = "qwen3:4b"
TOOL_FUNCTIONS: dict[str, Callable[..., dict[str, object]]] = {
list_files.__name__: list_files,
read_file.__name__: read_file,
run_tests.__name__: run_tests,
propose_patch.__name__: propose_patch,
}
SYSTEM_PROMPT = """
You are a careful local Python bug-fixing agent working only on sample_repo.
Paths passed to file tools are relative to sample_repo, so use "." for its root.
Use tools to inspect evidence instead of guessing. Prefer one tool call per reasoning step.
The test tool accepts only the scope "sample_repo" and is not a general shell.
When fixing a bug, run tests, read the relevant file, and propose the smallest exact replacement.
You can create a pending proposal, but you cannot apply it. Never claim that source changed.
After proposing a patch, tell the user to review it with the separate approval CLI.
""".strip()
def run_agent_turn(messages: list[object], user_input: str) -> None:
messages.append({"role": "user", "content": user_input})
while True:
response: ChatResponse = chat(
model=MODEL,
messages=messages,
tools=list(TOOL_FUNCTIONS.values()),
)
messages.append(response.message)
if response.message.tool_calls:
for tool_call in response.message.tool_calls:
tool_name = tool_call.function.name
function_to_call = TOOL_FUNCTIONS.get(tool_name)
if function_to_call is None:
result = {
"status": "error",
"error_message": f"Unknown tool: {tool_name}",
}
else:
print(f"Tool: {tool_name}({dict(tool_call.function.arguments)})")
try:
result = function_to_call(**tool_call.function.arguments)
except Exception as exc:
result = {"status": "error", "error_message": str(exc)}
print(f"Observation: {json.dumps(result, ensure_ascii=False)}")
messages.append(
{
"role": "tool",
"tool_name": tool_name,
"content": json.dumps(result, ensure_ascii=False),
}
)
continue
print(f"Agent: {response.message.content or '(No text response)'}")
return
def main() -> None:
messages: list[object] = [{"role": "system", "content": SYSTEM_PROMPT}]
print("Local bug agent ready. Type quit to exit.")
while True:
user_input = input("\nYou: ").strip()
if user_input.lower() in {"quit", "exit"}:
break
if user_input:
run_agent_turn(messages, user_input)
if __name__ == "__main__":
main()
This full-file reference exposes the proposal function beside the existing read plus test tools. The conversation loop remains unchanged.
Your agent can now record a minimal fix without applying it. Next up, you will review the unified diff and approve the change through a separate command that the model cannot access.
Review, Approve, and Prove the Fix
Your local agent loop has diagnosed the division bug and recorded a pending patch. The source still uses integer division because the model has no write tool.
A pending proposal must stay inert until a human approval process verifies its target. That process must also verify its content hash before changing the source.
In this step, get ready to:
- Build a separate approval CLI with stale-file protection.
- Review a unified diff before granting approval.
- Prove the reviewed fix through direct tests and the agent's restricted test tool.
Create the separate approval CLI
The approval CLI stays separate from TOOL_FUNCTIONS. Only you can start the process that performs the final write.
- In the Visual Studio Code Explorer from earlier, click the New File button in the Explorer toolbar.
- Enter approve_patch.py as the file name.
- Add the argument parser and target checks by pasting this first section into approve_patch.py.
from __future__ import annotations
import argparse
import difflib
import hashlib
import json
from patch_tool import PENDING_PATCH
from read_tools import WORKSPACE_ROOT, resolve_workspace_path
def main() -> None:
parser = argparse.ArgumentParser(
description="Preview or explicitly approve the pending agent patch."
)
parser.add_argument(
"--approve",
action="store_true",
help="Apply the reviewed patch after all checks pass.",
)
args = parser.parse_args()
if not PENDING_PATCH.is_file():
raise SystemExit("No pending patch exists.")
proposal = json.loads(PENDING_PATCH.read_text(encoding="utf-8"))
path = resolve_workspace_path(proposal["relative_path"])
if path.suffix != ".py" or not path.is_file():
raise SystemExit("Pending target is not an existing Python file.")
What Does This Section Do?
- The argument parser makes preview mode the default behavior.
- The --approve flag records your explicit decision to permit a write.
- PENDING_PATCH loads the inert proposal stored outside sample_repo.
- resolve_workspace_path() applies the existing workspace boundary to the proposed target.
- The suffix check limits approval to an existing .py file.
- Save approve_patch.py.
You should see approve_patch.py beside agent.py in the Explorer.
Approval File Missing?
Check that approve_patch.py sits beside agent.py. Check that its filename ends with .py.
Help me verify the location of approve_patch.py.
The next section creates the proposed replacement in memory. It also validates the original SHA-256 hash before any write can occur.
- In approve_patch.py, find the line containing Pending target is not an existing Python file..
- Paste this continuation directly below that line:
current = path.read_text(encoding="utf-8")
current_hash = hashlib.sha256(current.encode("utf-8")).hexdigest()
if current_hash != proposal["sha256"]:
raise SystemExit("The file changed after the proposal. Refusing stale patch.")
if current.count(proposal["old_text"]) != 1:
raise SystemExit("The expected old text no longer matches exactly once.")
updated = current.replace(proposal["old_text"], proposal["new_text"], 1)
relative_path = path.relative_to(WORKSPACE_ROOT).as_posix()
diff = difflib.unified_diff(
current.splitlines(keepends=True),
updated.splitlines(keepends=True),
fromfile=f"a/{relative_path}",
tofile=f"b/{relative_path}",
)
print(f"Rationale: {proposal['rationale']}")
print("".join(diff))
if not args.approve:
print("Preview only. No files changed. Re-run with --approve after review.")
return
path.write_text(updated, encoding="utf-8")
PENDING_PATCH.unlink()
print(f"Approved and applied: {relative_path}")
if __name__ == "__main__":
main()
How Is Approval Enforced?
- The stored hash must match the current source file.
- The expected old text must still occur exactly once.
- difflib.unified_diff() builds the reviewable change before the approval branch.
- path.write_text() runs only when --approve is present.
- PENDING_PATCH.unlink() removes the consumed proposal after a successful write.
- Save approve_patch.py.
Before you run the preview, how many source lines do you expect the diff to replace?
- Preview the pending patch by running this command:
python approve_patch.py
What Does the Preview Prove?
You should see the stored rationale followed by a diff for calculator.py. The diff removes return total // groups.
The diff adds return total / groups. The final message confirms that preview mode changed no files.
Preview Rejected the Proposal?
A missing proposal means .pending_patch.json is unavailable. A stale proposal means sample_repo/calculator.py changed after the model created the patch.
Help me diagnose the rejected pending proposal.
✔️ Awesome, I've got everything!
Your approval CLI now validates the target before showing a preview. It remains outside the agent's tool registry.
ⓧ I'd like to double check the full code
The complete approve_patch.py file is shown below for comparison.
from __future__ import annotations
import argparse
import difflib
import hashlib
import json
from patch_tool import PENDING_PATCH
from read_tools import WORKSPACE_ROOT, resolve_workspace_path
def main() -> None:
parser = argparse.ArgumentParser(
description="Preview or explicitly approve the pending agent patch."
)
parser.add_argument(
"--approve",
action="store_true",
help="Apply the reviewed patch after all checks pass.",
)
args = parser.parse_args()
if not PENDING_PATCH.is_file():
raise SystemExit("No pending patch exists.")
proposal = json.loads(PENDING_PATCH.read_text(encoding="utf-8"))
path = resolve_workspace_path(proposal["relative_path"])
if path.suffix != ".py" or not path.is_file():
raise SystemExit("Pending target is not an existing Python file.")
current = path.read_text(encoding="utf-8")
current_hash = hashlib.sha256(current.encode("utf-8")).hexdigest()
if current_hash != proposal["sha256"]:
raise SystemExit("The file changed after the proposal. Refusing stale patch.")
if current.count(proposal["old_text"]) != 1:
raise SystemExit("The expected old text no longer matches exactly once.")
updated = current.replace(proposal["old_text"], proposal["new_text"], 1)
relative_path = path.relative_to(WORKSPACE_ROOT).as_posix()
diff = difflib.unified_diff(
current.splitlines(keepends=True),
updated.splitlines(keepends=True),
fromfile=f"a/{relative_path}",
tofile=f"b/{relative_path}",
)
print(f"Rationale: {proposal['rationale']}")
print("".join(diff))
if not args.approve:
print("Preview only. No files changed. Re-run with --approve after review.")
return
path.write_text(updated, encoding="utf-8")
PENDING_PATCH.unlink()
print(f"Approved and applied: {relative_path}")
if __name__ == "__main__":
main()
This reference shows the complete preview path and the guarded approval branch in one file.
Review and approve the pending patch
The preview has verified that the proposal still targets the guarded workspace. It has also verified that the source matches the version the agent inspected.
Diff markers can feel fiddly on first read. Focus on the removed line and its replacement.
- Review the diff path for calculator.py.
- Confirm that the removed line is return total // groups.
- Confirm that the added line is return total / groups.
- Switch back to sample_repo/calculator.py in the editor.
You should still see return total // groups. Preview mode has preserved the original source.
The next command performs the only source write in this workflow. You have already reviewed the exact target and replacement.
- Apply the reviewed patch by running this command:
python approve_patch.py --approve
What Changes After Approval?
The CLI repeats its safety checks before writing the replacement. You should see Approved and applied: calculator.py after the write succeeds.
The consumed .pending_patch.json file is removed. The same proposal cannot be applied twice.
- Switch back to sample_repo/calculator.py.
- Refresh the Explorer in Visual Studio Code.
You should now see return total / groups in divide(). You should no longer see .pending_patch.json in the Explorer.
Approval Did Not Apply?
Confirm that the command includes --approve. If the proposal is stale, ask the agent to create a fresh proposal from the current source.
Help me troubleshoot the approval checks.
Prove the fix from both test paths
A successful write proves that approval worked. The Python unittest suite proves that the reviewed replacement fixed the original behavior.
Before you run the suite, do you expect the fractional case to agree with the other two tests?
- Run all three calculator tests by running this command:
python -m unittest discover -s sample_repo/tests -p "test_*.py"
What Do the Tests Prove?
You should see three tests complete with OK. The fractional case now returns 2.5.
The even-division case still passes. The zero-group guard still raises the expected error.
Still Seeing a Failed Test?
Check that sample_repo/calculator.py contains return total / groups. Check that your terminal is in the folder containing sample_repo.
Help me debug the calculator test failure.
The direct test run proves that the reviewed source works. The final check confirms that the agent's fixed test tool observes the same result.
- Return to the activated PowerShell window from earlier.
- Start the local agent by running this command:
python agent.py
What Should Start?
You should see Local bug agent ready. Type quit to exit. The agent still lacks access to approve_patch.py.
Before you ask the agent, do you expect its restricted test tool to agree with your direct test run?
- At the You: prompt, ask the agent to call run_tests with sample_repo.
- Ask the agent to summarize the returned test result.
What Does This Final Check Prove?
You should see tool activity for run_tests. The observation should report the passing suite.
The agent can verify the reviewed fix through its allowlisted command. It still cannot apply source changes.
- Close the agent loop by entering quit.
Agent Skipped the Test Tool?
Repeat the request with run_tests named explicitly. State that its required argument is sample_repo.
Help me prompt the agent to use its test tool.
That is the approval boundary proven. The model can investigate and propose while the final source write remains under your control.
Secret mission
Turn the Guardrails Into Regression Checks
Your guardrails blocked unsafe requests during the demo. Now you will turn those safety boundaries into executable regression checks that catch traversal, command injection, and out-of-workspace patch attempts.
Clean Up Your Resources
Clean Up Your Resources
Choose whether to keep your local project ready, pause its processes, or delete its files and model. This project has no paid cloud resources or API fees.
Resources you used:
- The local project folder containing agent.py, approve_patch.py, security_checks.py, tutorial-env, and sample_repo.
- The Ollama application installed on Windows.
- The locally stored Qwen3 4B model tagged as qwen3:4b.
Keep everything running
No action is needed. Choose this if you plan to demonstrate the agent again or keep developing its guardrails.
- Keep the local project folder with its passing tests and security checks.
- Leave the qwen3:4b model installed for future local inference.
- Leave Ollama running so the local agent can send model requests.
Pause - I'll come back to this later
Stop the running processes to free up memory. Your project files and local model remain available for your next session.
- Enter exit at the agent prompt to stop agent.py.
- Close Ollama in Windows to stop the local model server.
- Keep the project folder so the fixed calculator and guardrail checks remain on disk.
- Keep the qwen3:4b model so you do not need to download it again.
Delete - I don't want to use this again
Deleting these resources is permanent, so this option gives you a clean local reset. No paid cloud service is affected.
- Enter exit at the agent prompt if agent.py is still running.
- Press the Windows key to open search.
- Type File Explorer.
- Press Enter.
- Locate the project folder containing agent.py and sample_repo.
- Delete that project folder using File Explorer.
You should no longer see the project folder at its previous location.
- Press the Windows key to open search.
- Type PowerShell.
- Press Enter.
- Remove the local model by running this command:
ollama rm qwen3:4b
What does this command do?
The command removes the locally stored qwen3:4b model from Ollama. This releases about 2.5 GB of disk space.
- Check the PowerShell output after the command finishes.
You should see deleted 'qwen3:4b'.
Model Removal Not Working?
If PowerShell cannot run the command, confirm that Ollama is still installed. Retry the command before starting the application uninstall.
Help me troubleshoot the model cleanup.
- Press the Windows key to open search.
- Type Add or remove programs.
- Press Enter.
- Find Ollama in the installed applications list.
- Start Ollama's registered uninstaller.
- Follow the uninstall prompts until Windows confirms that Ollama has been removed.
Nice Work!
Nice Work!
You completed the full guardrailed workflow! Your local Python agent can investigate a bug while a separate approval process keeps the final source write under your control.
You've learned how to:
- Build a multi-turn agent loop with the Ollama Python SDK. Feed each tool observation back to Qwen3 4B before the next action.
- Enforce path containment for Python file access. Replace arbitrary shell access with one allowlisted unittest command.
- Create a hash-bound pending patch for a minimal fix. Keep source mutation behind a separate human approval CLI.
- Complete the optional Secret Mission by adding adversarial regression checks. These checks deny traversal reads. They reject injected test scopes. They block out-of-workspace patch proposals.
Ready to quiz yourself?