AI Retrieval Quality Gate
Build a TF-IDF retrieval API with evaluation metrics and CI regression tests.
Introduction
30 Second Summary
Search feels reliable until a familiar question is phrased in a different order. A result that looked perfect in one demo can miss the same idea moments later.
In this project, you will build a local technical knowledge search service with an interactive FastAPI interface. You will turn an exact-match baseline failure into a measured TF-IDF quality gate that GitHub Actions runs automatically.
What You'll Build
You enter a technical query in interactive API documentation to see retrieval-evaluation rank first before the same retrieval checks turn green on GitHub.
By the end of this project, you'll have:
- A technical search demo where each query returns ranked documents with cosine similarity scores.
- A measured comparison showing the exact-match baseline at 0.0 beside TF-IDF ranking at 1.0 for both Hit@1 and mean reciprocal rank.
- A browser-tested API with ranked JSON protected by a pytest regression suite that GitHub Actions runs on every push.
- Secret Mission: Add a confidence threshold that makes /search return an empty results list for unrelated queries.
Are there any prerequisites?
You need a Windows computer plus a GitHub account. No AI API key or paid model is required.
Before We Start
Before any hands-on work, lock in what this retrieval service does. This gives you a clear reason to measure ranking quality as you build.
Set Up the Windows Workspace
A retrieval gate is only trustworthy when your computer uses the same dependency versions as the automated quality checks. Package drift can hide failures locally before they surface in continuous integration.
In this step, you’ll confirm that Python, Git, and Visual Studio Code are ready on Windows.
You’ll also create a project-local virtual environment. That environment installs FastAPI, scikit-learn, and pytest at the pinned versions used throughout the project.
In this step, get ready to:
- Confirm that the required Windows tools are installed.
- Create the project folder with its dependency files.
- Build an isolated environment with every pinned dependency.
Prepare the Windows tools
Command Prompt gives you a direct way to check each system tool. Python Install Manager also keeps the required Python runtime separate from other versions on your computer.
- Press the Windows key to open the search bar.
- Type Command Prompt into the search bar.
- Press Enter to open Command Prompt.
- Check the installed Python runtimes by running:
py list
What does this command show?
The Python Install Manager lists every runtime available through the py command. Look for a Python 3.14 runtime in the output.
✔️ I see Python 3.14
That’s one foundation in place. Python 3.14 is available for the isolated environment you create later in this step.
ⓧ I see an older Python version
The manager is available. Your project still needs a Python 3.14 runtime.
- Install Python 3.14 by running these commands:
py install 3.14
py list
What do these commands do?
- The first command installs a CPython 3.14 runtime through Python Install Manager.
- The second command lists the available runtimes again.
You should now see Python 3.14 in the runtime list.
Python 3.14 missing from the list?
Check that the installation command completed before the runtime list appeared. If the installation stopped early, help me troubleshoot the Python runtime installation..
ⓧ Command not found
Python Install Manager needs to be installed before Command Prompt can provide the py command.
- Approve the installation if Windows displays a permission dialog.
- Install Python Install Manager by running:
winget install 9NQ7512CXL7T -e --accept-package-agreements --disable-interactivity
What does this command do?
Windows Package Manager installs Python Install Manager from its package identifier. The included options accept the package agreements while keeping the installation non-interactive.
Python Install Manager did not install?
Check that Command Prompt can access the internet. If the package installation still fails, help me troubleshoot the Python Install Manager command..
- Close Command Prompt after the installation finishes.
- Press the Windows key to open the search bar.
- Type Command Prompt into the search bar.
- Press Enter to reopen Command Prompt.
- Install the required runtime by running these commands:
py install 3.14
py list
What do these commands do?
- The first command installs the Python 3.14 runtime.
- The second command confirms that the runtime is available.
You should now see Python 3.14 in the runtime list.
The py command still unavailable?
Confirm that you reopened Command Prompt after installing the manager. If the command is still unavailable, help me find why the py command is missing..
- Check whether Git is installed by running:
git version
What does this command show?
The command prints the installed Git version. Any returned version confirms that Command Prompt can find Git.
✔️ I see a Git version
Git is ready to track the project files you create. You’ll use it when the retrieval quality gate moves to GitHub.
ⓧ Command not found
Git for Windows needs to be installed before Command Prompt can provide the git command.
- Approve the installation if Windows displays a permission dialog.
- Install Git for Windows by running:
winget install --id Git.Git -e --source winget
What does this command do?
Windows Package Manager installs the package with the exact Git.Git identifier. The source option retrieves it from the Windows Package Manager repository.
- Close Command Prompt after the installation finishes.
- Press the Windows key to open the search bar.
- Type Command Prompt into the search bar.
- Press Enter to reopen Command Prompt.
- Confirm the installation by running:
git version
What should I see?
You should see installed Git version information. That output confirms that the new Command Prompt session can find Git.
Git still unavailable?
Confirm that the Git installation completed before you reopened Command Prompt. If the command remains unavailable, help me troubleshoot Git for Windows..
Create the workspace files
Keeping the project on your Desktop gives it a predictable location. The folder also becomes the boundary for the virtual environment and every source file you add later.
- Return to Command Prompt from earlier.
- Create the ai-retrieval-quality-gate folder on your Desktop by running these commands:
cd /d "%USERPROFILE%\Desktop"
mkdir ai-retrieval-quality-gate
cd ai-retrieval-quality-gate
What do these commands do?
- The first command moves Command Prompt to your Desktop.
- The second command creates the ai-retrieval-quality-gate folder.
- The final command moves Command Prompt into that folder.
Your Command Prompt path should now end with Desktop\ai-retrieval-quality-gate.
Folder command failed?
A folder with the same name may already exist on your Desktop. If Command Prompt reports a folder problem, help me diagnose the workspace path..
- Open the current project folder in Visual Studio Code by running:
code .
What does this command do?
The code command starts Visual Studio Code. The dot tells the editor to open the current ai-retrieval-quality-gate folder.
✔️ Visual Studio Code opens
Your editor is connected to the correct folder. The Explorer sidebar should show ai-retrieval-quality-gate as the open workspace.
ⓧ Command not found
Visual Studio Code needs to be installed with command-line access. The Windows User setup adds the code command to your path.
- Use the official Windows setup guide to download the Visual Studio Code User setup installer.
- Complete the Visual Studio Code installation.
- Close Command Prompt after the installation finishes.
- Press the Windows key to open the search bar.
- Type Command Prompt into the search bar.
- Press Enter to reopen Command Prompt.
- Return to the workspace and open it by running these commands:
cd /d "%USERPROFILE%\Desktop\ai-retrieval-quality-gate"
code .
What do these commands do?
- The first command returns Command Prompt to the project folder.
- The second command opens that folder in Visual Studio Code.
Visual Studio Code still not opening?
Confirm that you installed the Windows User setup. Confirm that you restarted Command Prompt afterward. If the command still fails, help me troubleshoot the code command..
The dependency file makes this workspace reproducible. Each installation uses the same tested package versions.
- Use the Explorer sidebar’s file control to create requirements.txt inside ai-retrieval-quality-gate.
- Paste these dependency pins into requirements.txt:
fastapi[standard]==0.142.2
scikit-learn==1.9.1
pytest==9.1.1
What do these pins control?
- The FastAPI pin provides the web framework plus its standard development tools.
- The scikit-learn pin provides the ranking utilities used by the retriever.
- The pytest pin provides the test runner used by the release gate.
- Save requirements.txt.
- Confirm that requirements.txt appears in the Explorer sidebar.
Dependency file missing?
Check that the file sits directly inside ai-retrieval-quality-gate. If the file name or location looks different, help me fix my requirements file..
Generated caches and installed packages do not belong in source control. The ignore file keeps those disposable files outside future commits.
- Use the Explorer sidebar’s file control to create .gitignore inside ai-retrieval-quality-gate.
- Paste these ignore rules into .gitignore:
.venv/
__pycache__/
.pytest_cache/
What do these rules exclude?
- The .venv/ rule excludes the local virtual environment.
- The __pycache__/ rule excludes Python bytecode caches.
- The .pytest_cache/ rule excludes pytest’s local cache.
- Save .gitignore.
- Confirm that .gitignore appears beside requirements.txt in the Explorer sidebar.
Ignore file named incorrectly?
Confirm that the file name begins with a dot. Remove any hidden .txt ending. If the file still looks wrong, help me correct my .gitignore file..
✔️ Awesome, I've got everything!
Great. Double-check that both files are saved inside ai-retrieval-quality-gate.
ⓧ I'd like to double check the full code
Compare each file with the complete versions below.
fastapi[standard]==0.142.2
scikit-learn==1.9.1
pytest==9.1.1
What should match?
Your requirements.txt file should contain these three dependency pins in this order.
.venv/
__pycache__/
.pytest_cache/
What should match?
Your .gitignore file should contain these three directory rules in this order.
Build the isolated environment
A project-local environment keeps these dependencies away from other Python projects on your computer. Activating it also directs installation commands into .venv.
- Return to Command Prompt from earlier.
- Create the Python 3.14 virtual environment by running:
py -V:3.14 -m venv .venv
What does this command do?
The runtime selector chooses Python 3.14. The virtual environment module creates an isolated environment in the .venv directory.
- Confirm that .venv appears in the Visual Studio Code Explorer sidebar.
Virtual environment missing?
Confirm that Command Prompt is inside ai-retrieval-quality-gate. If Python reports a runtime problem, help me troubleshoot the virtual environment command..
- Activate the virtual environment in Command Prompt by running:
.venv\Scripts\activate
What does activation change?
Activation places the environment’s executable directory at the front of the current Command Prompt session. Python package commands now target .venv.
You should see (.venv) at the start of the Command Prompt line.
Activation marker missing?
Confirm that the .venv directory exists in the project folder. If the prompt still lacks the activation marker, help me troubleshoot Windows virtual environment activation..
- Install the pinned project dependencies by running:
py -m pip install -r requirements.txt
What does this command install?
The package installer reads every pin from requirements.txt. It installs those exact versions into the activated environment.
When installation finishes, Command Prompt returns to a line beginning with (.venv).
Dependency installation failed?
Confirm that requirements.txt is saved in the current folder. Confirm that the prompt begins with (.venv). If the installation still fails, help me diagnose the dependency installation..
Before you run the final check, what message do you expect when all three imports succeed?
- Verify the installed dependencies by running:
python -c "import fastapi, sklearn, pytest; print('Workspace ready')"
What should I see?
You should see Workspace ready on its own line. That message proves that FastAPI, scikit-learn, and pytest can all be imported from the activated environment.
Import check did not pass?
Check that Command Prompt still begins with (.venv). If the command reports a missing dependency, run the requirements installation again. For further help, help me debug the workspace verification..
Your Windows workspace now has the tools, pinned dependencies, and isolated environment needed for repeatable retrieval experiments. Next up, you’ll build an exact-match retriever and watch it fail on a reordered query.
Build a Baseline Retriever
Your Windows workspace can now run the project's pinned dependencies. This step gives that environment its first retrieval behavior.
An exact-match baseline accepts a query only when the whole normalized phrase appears in a document. A copied phrase can succeed while the same ideas in a new order fail.
In this step, get ready to:
- Build a shared collection of technical knowledge documents.
- Implement an exact-match search function.
- Run copied and reordered queries to expose the baseline's limitation.
Create the knowledge corpus
A retrieval corpus is the collection of information your search function examines. Each document in data.py has a stable ID plus searchable technical text.
- In the Visual Studio Code file sidebar, select the ai-retrieval-quality-gate folder.
- Use the new-file icon at the top of the file sidebar to create data.py.
- Add the first two knowledge entries to data.py by pasting this code:
DOCUMENTS = [
{
"id": "api-testing",
"title": "API Testing",
"text": (
"Use automated tests to verify endpoint status codes, response schemas, "
"and error handling before deployment."
),
},
{
"id": "retrieval-evaluation",
"title": "Retrieval Evaluation",
"text": (
"Evaluate search quality with labeled queries and ranking metrics such "
"as hit rate and reciprocal rank."
),
},
]
What does this data represent?
- The id value gives each document a stable identifier that search results can return.
- The title value gives each knowledge entry a readable name.
- The text value holds the technical content that the retriever searches.
- Save data.py.
- Check the file sidebar for data.py.
You should see data.py inside ai-retrieval-quality-gate. The editor should show entries for api-testing and retrieval-evaluation.
Does the document list look incomplete?
Check that DOCUMENTS starts with an opening square bracket. Confirm that the final line contains the matching closing square bracket.
Need help checking the structure? Help me find a bracket or comma problem in my Python DOCUMENTS list.
The remaining entries cover monitoring plus continuous integration. Adding them gives the baseline four distinct topics to search.
- In data.py, place your cursor immediately above the final closing square bracket.
- Add the remaining knowledge entries by pasting this code:
{
"id": "model-monitoring",
"title": "Model Monitoring",
"text": (
"Track latency, failures, drift, and quality regressions after an AI "
"system reaches production."
),
},
{
"id": "ci-pipelines",
"title": "Continuous Integration",
"text": (
"Run tests automatically on every code change so regressions are caught "
"before merging."
),
},
Why these two entries?
The monitoring entry contains the phrase used by the exact-query demonstration. Its words also appear in a different order in the second demonstration query.
The continuous integration entry gives the collection another technical topic. Distinct topics make incorrect retrieval easier to spot.
- Save data.py.
- Count the id entries in the editor.
You should see four document IDs. The final two should be model-monitoring and ci-pipelines.
Are the new entries outside the list?
Move both new document objects above the closing square bracket for DOCUMENTS. Keep the comma after each document object.
Still unsure about the nesting? Help me place two document dictionaries inside my existing Python list.
Implement the exact-match search
The baseline normalizes letter case before checking whether the complete query appears in each document. This simple containment rule makes its behavior deterministic.
- In the VS Code file sidebar, select the ai-retrieval-quality-gate folder.
- Use the new-file icon at the top of the file sidebar to create retriever.py.
- Implement the baseline functions in retriever.py by pasting this code:
from data import DOCUMENTS
def document_text(document):
return f"{document['title']} {document['text']}"
def baseline_search(query):
normalized_query = query.strip().lower()
if not normalized_query:
return []
for document in DOCUMENTS:
if normalized_query in document_text(document).lower():
return [{**document, "score": 1.0}]
return []
What does this code do?
- The document_text() function combines a document's title with its main text.
- The baseline_search() function removes surrounding spaces from the query. It also converts the query to lowercase.
- The containment check returns the first document containing the complete normalized query.
- An empty query or missing match returns an empty list.
- Save retriever.py.
- Check the file sidebar for retriever.py.
You should see retriever.py beside data.py. The editor should contain document_text() plus baseline_search().
Is the import highlighted as a problem?
Confirm that both files are directly inside ai-retrieval-quality-gate. Check that the data file is named data.py with lowercase letters.
Need help tracing the import? Help me check why retriever.py cannot import DOCUMENTS from data.py.
The demonstration block sends two queries through the same baseline. Their different outcomes reveal whether matching the same ideas requires the same word order.
- In retriever.py, place your cursor below the final return [] line.
- Add the demonstration block by pasting this code:
if __name__ == "__main__":
exact_query = "Track latency, failures, drift"
reordered_query = "quality regressions drift latency production"
exact_result = baseline_search(exact_query)
reordered_result = baseline_search(reordered_query)
print("Exact query:", exact_result[0]["id"] if exact_result else "NO MATCH")
print(
"Reordered query:",
reordered_result[0]["id"] if reordered_result else "NO MATCH",
)
How does the demonstration work?
- The exact query copies a complete phrase from the monitoring document.
- The reordered query uses relevant monitoring words in a different sequence.
- Each print statement displays the matching document ID. It displays NO MATCH when the returned list is empty.
- Save retriever.py.
Before you run the script, which query do you think survives the complete-string containment check?
- Test the baseline in the activated Command Prompt by running this command:
python retriever.py
What should you see?
You should first see Exact query: model-monitoring. You should next see Reordered query: NO MATCH.
The reordered query contains the important monitoring words. Their new order prevents the complete string from appearing in the document.
Did the script stop before printing both results?
Confirm that your Command Prompt is still inside the ai-retrieval-quality-gate folder. Check both Python files for missing commas or closing brackets.
Still stuck? Help me diagnose why python retriever.py does not print both baseline results.
✔️ Awesome, I've got everything!
Your retriever.py now matches the complete baseline implementation.
ⓧ I'd like to double check the full code
from data import DOCUMENTS
def document_text(document):
return f"{document['title']} {document['text']}"
def baseline_search(query):
normalized_query = query.strip().lower()
if not normalized_query:
return []
for document in DOCUMENTS:
if normalized_query in document_text(document).lower():
return [{**document, "score": 1.0}]
return []
if __name__ == "__main__":
exact_query = "Track latency, failures, drift"
reordered_query = "quality regressions drift latency production"
exact_result = baseline_search(exact_query)
reordered_result = baseline_search(reordered_query)
print("Exact query:", exact_result[0]["id"] if exact_result else "NO MATCH")
print(
"Reordered query:",
reordered_result[0]["id"] if reordered_result else "NO MATCH",
)
Add the labeled evaluation cases
A labeled evaluation set pairs each query with the document ID that should rank first. These labels make retrieval comparisons reproducible.
- Switch back to data.py in VS Code.
- Place your cursor below the closing square bracket for DOCUMENTS.
- Add the four labeled cases by pasting this code:
EVAL_CASES = [
{
"query": "response schemas endpoint status codes",
"expected_id": "api-testing",
},
{
"query": "ranking metrics labeled queries search quality",
"expected_id": "retrieval-evaluation",
},
{
"query": "quality regressions drift latency production",
"expected_id": "model-monitoring",
},
{
"query": "code change tests automatically continuous integration",
"expected_id": "ci-pipelines",
},
]
What do these cases capture?
- Each query value contains the wording that the retriever evaluates.
- Each expected_id value identifies the correct top result.
- The four cases cover every document in the corpus once.
- Save data.py.
Before you rerun the baseline, do you expect the added evaluation cases to change its two demonstration results?
- Verify the complete baseline from the activated Command Prompt by running this command:
python retriever.py
What does the final check prove?
You should still see Exact query: model-monitoring followed by Reordered query: NO MATCH.
The successful run confirms that data.py remains valid after adding the labels. The visible miss preserves the limitation this baseline is designed to demonstrate.
Did the final check break after adding the cases?
Confirm that EVAL_CASES starts below the closing bracket for DOCUMENTS. Check that every case ends with a comma.
Need another pair of eyes? Help me check the EVAL_CASES structure in data.py.
✔️ Awesome, I've got everything!
Your data.py now contains all four documents plus all four labeled evaluation cases.
ⓧ I'd like to double check the full code
DOCUMENTS = [
{
"id": "api-testing",
"title": "API Testing",
"text": (
"Use automated tests to verify endpoint status codes, response schemas, "
"and error handling before deployment."
),
},
{
"id": "retrieval-evaluation",
"title": "Retrieval Evaluation",
"text": (
"Evaluate search quality with labeled queries and ranking metrics such "
"as hit rate and reciprocal rank."
),
},
{
"id": "model-monitoring",
"title": "Model Monitoring",
"text": (
"Track latency, failures, drift, and quality regressions after an AI "
"system reaches production."
),
},
{
"id": "ci-pipelines",
"title": "Continuous Integration",
"text": (
"Run tests automatically on every code change so regressions are caught "
"before merging."
),
},
]
EVAL_CASES = [
{
"query": "response schemas endpoint status codes",
"expected_id": "api-testing",
},
{
"query": "ranking metrics labeled queries search quality",
"expected_id": "retrieval-evaluation",
},
{
"query": "quality regressions drift latency production",
"expected_id": "model-monitoring",
},
{
"query": "code change tests automatically continuous integration",
"expected_id": "ci-pipelines",
},
]
You have reproduced the baseline's weakness with a result you can see. Next up, you'll replace complete-string matching with TF-IDF ranking plus cosine similarity. You will measure the improvement across the same four labeled queries.
Measure TF-IDF Ranking
Your exact-match retriever reproduced the weakness you needed to see. A reordered query missed the correct document even though it carried the same ideas.
Now you'll use TF-IDF with cosine similarity to rank documents by relevance. You'll measure both retrievers with Hit@1 plus mean reciprocal rank.
In this step, get ready to:
- Fit a TF-IDF index over the four knowledge documents.
- Rank queries with cosine similarity.
- Compare the baseline with the improved retriever using Hit@1 and MRR.
Add TF-IDF ranking
The scikit-learn library converts each document into weighted term features. Terms that help distinguish one document from the others receive more influence.
Why TF-IDF for this gate?
TF-IDF gives this project deterministic rankings without a paid API. The same query against the same documents produces the same result.
A local language model adds model downloads plus variable output. TF-IDF keeps your focus on evaluation quality.
- In retriever.py, find this import at the top of the file:
from data import DOCUMENTS
What is already imported?
This import gives the retriever access to the four documents in data.py. The ranking components need the same collection.
- Replace that one-line import with this import group:
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity
from data import DOCUMENTS
What changed?
- TfidfVectorizer converts document text into weighted numerical features.
- cosine_similarity compares the query vector with every document vector.
- The existing DOCUMENTS import continues to supply the searchable content.
The retriever needs a fitted document matrix before it can compare a query. This matrix becomes the reusable index for every search.
- Add the ranking setup directly above if __name__ == "__main__": by pasting this code:
vectorizer = TfidfVectorizer(ngram_range=(1, 2))
document_matrix = vectorizer.fit_transform(
[document_text(document) for document in DOCUMENTS]
)
What does this setup do?
- The vectorizer learns features from individual words plus two-word phrases.
- The list comprehension combines each document title with its text.
- The fitted document_matrix stores a numerical representation of all four documents.
- Save retriever.py.
- Check that the fitted matrix loads successfully by running:
python retriever.py
What does this check prove?
The script fits the document matrix before reaching its demonstration block. Seeing the two baseline lines proves that the new imports plus the fitted index load without stopping the script.
You'll still see model-monitoring for the exact query. You'll still see NO MATCH for the reordered query.
Does the script stop before printing?
Confirm that the two scikit-learn imports sit above the existing DOCUMENTS import. Check that your activated terminal still shows (.venv) before its prompt.
Still stuck? Help me troubleshoot why retriever.py stops while creating its TF-IDF document matrix.
A query must pass through the fitted vocabulary before cosine similarity can compare it with the document matrix. The resulting scores provide the ranking signal that exact substring matching lacks.
- Add search(query, top_k=3) directly below the document_matrix assignment by pasting this code:
def search(query, top_k=3):
normalized_query = query.strip()
if not normalized_query:
return []
query_vector = vectorizer.transform([normalized_query])
scores = cosine_similarity(query_vector, document_matrix)[0]
ranked_indices = sorted(
range(len(DOCUMENTS)),
key=lambda index: scores[index],
reverse=True,
)[:top_k]
return [
{
**DOCUMENTS[index],
"score": round(float(scores[index]), 4),
}
for index in ranked_indices
]
How does the ranked search work?
- The empty-query check returns an empty list when the input contains only whitespace.
- The fitted vectorizer transforms the query into the same feature space as the documents.
- Cosine similarity produces one relevance score for each document.
- The sorted indices place the strongest score first before limiting the results with top_k.
- Save retriever.py.
- Confirm that the complete module still runs by executing:
python retriever.py
Why does the baseline output remain?
The demonstration block still calls baseline_search(). Its familiar output confirms that adding search() preserved the original baseline for comparison.
Your retriever now contains both systems in one file. The original failure remains reproducible while the ranked alternative is ready for evaluation.
Does retriever.py fail after the new function?
Check that search() begins at the left edge of the file. Confirm that its return list stays inside the function.
Need another pair of eyes? Help me compare my search function with the expected indentation and ranking flow.
✔️ Awesome, I've got everything!
Great. Confirm that retriever.py is saved before you build the evaluation harness.
ⓧ I'd like to double check the full code
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity
from data import DOCUMENTS
def document_text(document):
return f"{document['title']} {document['text']}"
def baseline_search(query):
normalized_query = query.strip().lower()
if not normalized_query:
return []
for document in DOCUMENTS:
if normalized_query in document_text(document).lower():
return [{**document, "score": 1.0}]
return []
vectorizer = TfidfVectorizer(ngram_range=(1, 2))
document_matrix = vectorizer.fit_transform(
[document_text(document) for document in DOCUMENTS]
)
def search(query, top_k=3):
normalized_query = query.strip()
if not normalized_query:
return []
query_vector = vectorizer.transform([normalized_query])
scores = cosine_similarity(query_vector, document_matrix)[0]
ranked_indices = sorted(
range(len(DOCUMENTS)),
key=lambda index: scores[index],
reverse=True,
)[:top_k]
return [
{
**DOCUMENTS[index],
"score": round(float(scores[index]), 4),
}
for index in ranked_indices
]
if __name__ == "__main__":
exact_query = "Track latency, failures, drift"
reordered_query = "quality regressions drift latency production"
exact_result = baseline_search(exact_query)
reordered_result = baseline_search(reordered_query)
print("Exact query:", exact_result[0]["id"] if exact_result else "NO MATCH")
print(
"Reordered query:",
reordered_result[0]["id"] if reordered_result else "NO MATCH",
)
How to use this reference
Compare this complete file with your saved retriever.py. Pay close attention to the import order plus the placement of the fitted matrix.
Build the evaluation harness
A labeled evaluation set pairs each query with the document that should rank first. Running every retriever against the same cases makes the comparison reproducible.
Reciprocal rank assigns a score based on the first correct result's position. A missing expected result receives 0.0.
- In the VS Code Explorer sidebar, create evaluate.py inside the ai-retrieval-quality-gate folder.
- Add the evaluation imports plus reciprocal_rank() by pasting this code:
from data import EVAL_CASES
from retriever import baseline_search, search
def reciprocal_rank(results, expected_id):
for rank, result in enumerate(results, start=1):
if result["id"] == expected_id:
return 1 / rank
return 0.0
What does reciprocal rank measure?
- The imports bring in the four labeled cases plus both retrieval functions.
- The loop starts ranking at one so a first-place match receives a reciprocal rank of one.
- The final return gives missing expected documents a score of zero.
- Save evaluate.py.
You'll see evaluate.py beside data.py plus retriever.py in the Explorer sidebar.
Can't find evaluate.py?
Check that the file appears directly inside ai-retrieval-quality-gate. A file created inside .venv cannot import the project modules from the expected location.
Need help locating it? Help me verify that evaluate.py is saved in the correct VS Code workspace folder.
The evaluator needs to apply one retrieval function to every labeled case. It then averages first-place success plus reciprocal rank across the full set.
- Add evaluate() below reciprocal_rank() by pasting this code:
def evaluate(search_function):
hit_at_1_scores = []
reciprocal_ranks = []
for case in EVAL_CASES:
results = search_function(case["query"])
hit_at_1_scores.append(
float(bool(results) and results[0]["id"] == case["expected_id"])
)
reciprocal_ranks.append(reciprocal_rank(results, case["expected_id"]))
return {
"hit_at_1": round(sum(hit_at_1_scores) / len(hit_at_1_scores), 2),
"mrr": round(sum(reciprocal_ranks) / len(reciprocal_ranks), 2),
}
if __name__ == "__main__":
print("Baseline:", evaluate(baseline_search))
print("TF-IDF:", evaluate(search))
How does the evaluator compare systems?
- The evaluator accepts a retrieval function so both systems receive the same test cases.
- The Hit@1 score records whether the expected document occupies the first position.
- The reciprocal-rank list records how early the expected document appears.
- The main block prints one report for the baseline plus one report for TF-IDF.
- Save evaluate.py.
Run and compare the evaluation
The baseline plus TF-IDF now face the same four labeled queries. Their metric values reveal whether ranked retrieval fixes the exact-match failure across the complete test set.
Before you run the evaluation, which retriever do you expect to rank every correct document first?
- Measure both retrievers by running:
python evaluate.py
What does this command measure?
The command loads all four EVAL_CASES from data.py. It passes each query through both retrieval functions.
The resulting report compares the two systems under identical conditions. That shared evaluation set turns an improvement claim into measurable evidence.
You'll see Baseline: {'hit_at_1': 0.0, 'mrr': 0.0} first. Every labeled query uses reordered language that the exact-substring baseline misses.
You'll then see TF-IDF: {'hit_at_1': 1.0, 'mrr': 1.0}. All four expected documents rank first.
You now have evidence that the ranked retriever fixes the baseline failure across the whole labeled set. That measurement is the foundation of your retrieval quality gate.
How to read the metrics
Hit@1 measures the share of queries whose expected document ranks first. A value of 1.0 means all four cases succeeded.
MRR averages the reciprocal rank of the first correct result. Its value also reaches 1.0 because every expected document occupies rank one.
Are the metrics missing or different?
Confirm that evaluate.py imports search from retriever.py. Check that evaluate(search) appears in the final print statement.
If TF-IDF scores stay below 1.0, compare the ngram_range value plus the descending score sort with the reference code.
Still seeing unexpected values? Help me trace why my evaluation metrics differ from the expected baseline and TF-IDF results.
✔️ Awesome, I've got everything!
Excellent. Save evaluate.py with the report still visible in your terminal.
ⓧ I'd like to double check the full code
from data import EVAL_CASES
from retriever import baseline_search, search
def reciprocal_rank(results, expected_id):
for rank, result in enumerate(results, start=1):
if result["id"] == expected_id:
return 1 / rank
return 0.0
def evaluate(search_function):
hit_at_1_scores = []
reciprocal_ranks = []
for case in EVAL_CASES:
results = search_function(case["query"])
hit_at_1_scores.append(
float(bool(results) and results[0]["id"] == case["expected_id"])
)
reciprocal_ranks.append(reciprocal_rank(results, case["expected_id"]))
return {
"hit_at_1": round(sum(hit_at_1_scores) / len(hit_at_1_scores), 2),
"mrr": round(sum(reciprocal_ranks) / len(reciprocal_ranks), 2),
}
if __name__ == "__main__":
print("Baseline:", evaluate(baseline_search))
print("TF-IDF:", evaluate(search))
How to use this reference
Compare this file with your saved evaluate.py. Check the metric calculations plus the two function calls in the main block.
Your retriever now has a measured quality improvement instead of a hand-picked success. Next, you'll expose that ranked search through a documented FastAPI interface.
Serve the Retriever API
Your retriever now ranks all four labeled queries correctly. The results currently live inside local Python scripts.
A measured component becomes useful to other software through a stable interface. That interface needs to validate inputs.
In this step, you'll wrap search() in FastAPI. Your endpoints will return predictable JSON.
In this step, get ready to:
- Create the FastAPI application.
- Launch the interactive documentation from the development server.
- Test the health endpoint plus the ranked search endpoint.
Build the API interface
An endpoint gives other software a defined place to send a request. Each endpoint calls a Python function that returns structured data.
- In the VS Code Explorer sidebar, create main.py inside the ai-retrieval-quality-gate folder.
- Add the API application by copying this code into main.py:
from fastapi import FastAPI
from retriever import search
app = FastAPI(title="AI Retrieval Quality Gate", version="1.0.0")
@app.get("/health")
def health():
return {"status": "ok"}
@app.get("/search")
def retrieve(q: str, top_k: int = 3):
return {"query": q, "results": search(q, top_k)}
What Does This API Code Do?
- FastAPI creates the application stored in app.
- The title identifies the service in its interactive documentation.
- The version is set to 1.0.0.
- The /health endpoint provides a simple service status.
- The /search endpoint exposes your existing retriever.
- FastAPI interprets non-path function parameters as query parameters.
- The q parameter is required.
- The top_k parameter is optional with a default value of 3.
- The retrieve() function passes both values to search().
- Save main.py.
- Confirm main.py appears in the VS Code Explorer sidebar.
✔️ Awesome, I've got everything!
Your main.py file now contains the application plus both endpoint functions.
ⓧ I'd like to double check the full code
- Compare your saved main.py with the reference below.
from fastapi import FastAPI
from retriever import search
app = FastAPI(title="AI Retrieval Quality Gate", version="1.0.0")
@app.get("/health")
def health():
return {"status": "ok"}
@app.get("/search")
def retrieve(q: str, top_k: int = 3):
return {"query": q, "results": search(q, top_k)}
What Should Match?
The application reuses search() from retriever.py. One ranking implementation now serves both the evaluation path and the API path.
Launch the interactive documentation
The development server imports app from main.py. It then makes your endpoints available through a local web address.
- Launch the application from the activated terminal by running:
fastapi dev
What Does This Command Do?
- The fastapi dev command detects main.py in your current folder.
- It starts the development server for the FastAPI application.
- It makes the interactive documentation available at http://127.0.0.1:8000/docs.
Your terminal confirms that the server is running. Keep this terminal active while you test the endpoints.
- Load the interactive documentation by entering http://127.0.0.1:8000/docs in your browser's address bar.
You'll see a documentation page containing /health plus /search. The page provides controls for sending requests to both endpoints.
Development Server Not Starting?
- Confirm the terminal prompt still shows your activated .venv environment.
- Check that main.py is saved inside ai-retrieval-quality-gate.
- Check that the application variable is spelled app.
Still stuck? Help me diagnose why my FastAPI development server is not starting.
Test both API endpoints
The interactive documentation sends real requests to your running application. Its response panel shows the exact data another program would receive.
- Expand the /health endpoint.
- Enable its request controls.
- Submit the health request.
You'll see {"status":"ok"} in the response body. This confirms that the application is accepting requests.
- Expand the /search endpoint.
- Enable its request controls.
- Enter ranking metrics labeled queries search quality for q.
- Keep 3 for top_k.
Before you submit the search request, take a moment to predict the first-ranked document.
- Submit the search request.
The first item in results has the ID retrieval-evaluation. This proves the API preserves the ranking measured in your evaluation.
Each result includes an id. Each result also includes a title.
You can also see the document text. The score shows its similarity value.
You did it. Your measured retriever now answers browser-based API requests with ranked results.
Your retrieval API is live on your computer. Next, you'll turn its evaluation metrics into automated tests that protect every future change.
Create the Release Gate
Your FastAPI service already ranks technical documents through a browser-tested endpoint. The measured results show that its ranking works today.
A metric run only by hand cannot stop a future change from silently degrading retrieval quality. This step converts your evaluation into continuous integration with GitHub Actions.
In this step, get ready to:
- Turn retrieval quality into repeatable tests.
- Define the GitHub Actions workflow that runs those tests.
- Publish the project through a green quality gate.
Add regression tests
A regression test turns a working behavior into a rule that future changes must preserve. The pytest framework checks those rules automatically.
- Switch back to the terminal running the development server.
- Press Ctrl+C to stop the development server.
- In the VS Code file sidebar, select the folder-creation control.
- Enter tests as the folder name.
- Select the tests folder in the file sidebar.
- Select the file-creation control.
- Enter test_retrieval.py as the file name.
- Add the quality checks by copying this code into tests/test_retrieval.py:
from evaluate import evaluate
from retriever import baseline_search, search
def test_tfidf_beats_baseline_and_meets_gate():
baseline_metrics = evaluate(baseline_search)
improved_metrics = evaluate(search)
assert improved_metrics["hit_at_1"] > baseline_metrics["hit_at_1"]
assert improved_metrics["mrr"] > baseline_metrics["mrr"]
assert improved_metrics["hit_at_1"] >= 1.0
assert improved_metrics["mrr"] >= 1.0
def test_empty_query_returns_no_results():
assert search(" ") == []
What do these tests protect?
- The first test evaluates the baseline retriever against the improved retriever.
- Its comparison proves that the improved retriever performs better.
- Its thresholds require perfect Hit@1 and MRR on the four labeled cases.
- The second test confirms that whitespace-only input returns no results.
- Save tests/test_retrieval.py.
Before you run the tests, do you expect both retrieval rules to pass?
- Run the local quality gate in the activated terminal with this command:
pytest
What does this command check?
The command discovers the test functions under tests. It reports whether each retrieval rule passes.
You should see that two tests passed. Good work. Your local gate now protects the behaviors your service depends on.
Tests not passing?
- Confirm that test_retrieval.py is inside the tests folder.
- Check each imported function name against evaluate.py and retriever.py.
- Confirm that the terminal still shows the activated .venv environment.
Help me debug my local retrieval tests.
✔️ Awesome, I've got everything!
Your local test file is saved. Both retrieval tests pass.
ⓧ I'd like to double check the full code
from evaluate import evaluate
from retriever import baseline_search, search
def test_tfidf_beats_baseline_and_meets_gate():
baseline_metrics = evaluate(baseline_search)
improved_metrics = evaluate(search)
assert improved_metrics["hit_at_1"] > baseline_metrics["hit_at_1"]
assert improved_metrics["mrr"] > baseline_metrics["mrr"]
assert improved_metrics["hit_at_1"] >= 1.0
assert improved_metrics["mrr"] >= 1.0
def test_empty_query_returns_no_results():
assert search(" ") == []
Add workflow and project documentation
Local tests only help when someone remembers to run them. The workflow runs the same quality gate for every configured repository event.
- In the VS Code file sidebar, select the folder-creation control.
- Enter .github as the folder name.
- Select the .github folder.
- Select the folder-creation control again.
- Enter workflows as the folder name.
You should now see the .github/workflows path in the file sidebar.
- Select the workflows folder.
- Create a file named evals.yml with the file-creation control.
- Define the automated quality gate by copying this workflow into .github/workflows/evals.yml:
name: Retrieval Evals
on:
push:
pull_request:
permissions:
contents: read
jobs:
test:
runs-on: ubuntu-latest
steps:
- name: Check out repository
uses: actions/checkout@v7.0.1
- name: Set up Python
uses: actions/setup-python@v7.0.0
with:
python-version: "3.14.8"
cache: "pip"
- name: Install dependencies
run: pip install -r requirements.txt
- name: Run retrieval quality gate
run: pytest
How does the workflow enforce quality?
- The push trigger starts the workflow after pushed changes.
- The pull_request trigger checks proposed changes before merging.
- The workflow gives its token read-only access to repository contents.
- The runner checks out the repository on ubuntu-latest.
- The setup action installs Python 3.14.8 with dependency caching.
- The final step runs the same pytest gate you tested locally.
- Save .github/workflows/evals.yml.
- Confirm that the file sidebar lists evals.yml inside .github/workflows.
Workflow file in the wrong place?
- Check that the complete path is .github/workflows/evals.yml.
- Confirm that the filename ends with .yml.
- Match the YAML indentation shown above with spaces.
Help me check my GitHub Actions workflow structure.
✔️ Awesome, I've got everything!
Your workflow file is saved in the directory GitHub Actions monitors.
ⓧ I'd like to double check the full code
name: Retrieval Evals
on:
push:
pull_request:
permissions:
contents: read
jobs:
test:
runs-on: ubuntu-latest
steps:
- name: Check out repository
uses: actions/checkout@v7.0.1
- name: Set up Python
uses: actions/setup-python@v7.0.0
with:
python-version: "3.14.8"
cache: "pip"
- name: Install dependencies
run: pip install -r requirements.txt
- name: Run retrieval quality gate
run: pytest
A clear project guide makes the measured result reproducible. Your README records the quality comparison plus every command needed to run the service.
- Select the ai-retrieval-quality-gate folder in the VS Code file sidebar.
- Create a file named README.md with the file-creation control.
- Document the project by copying this content into README.md:
# AI Retrieval Quality Gate
A local retrieval API that compares a brittle exact-match baseline with TF-IDF ranking and blocks quality regressions with automated tests.
## Measured result
| Retriever | Hit@1 | MRR |
| --- | ---: | ---: |
| Exact-match baseline | 0.0 | 0.0 |
| TF-IDF ranking | 1.0 | 1.0 |
## Run on Windows
```text
py -V:3.14 -m venv .venv
.venv\Scripts\activate
py -m pip install -r requirements.txt
python evaluate.py
pytest
fastapi dev
```
Open `http://127.0.0.1:8000/docs` to test the API.
What does the README capture?
- The opening explains the exact-match comparison plus the automated quality gate.
- The table records the measured Hit@1 and MRR results.
- The Windows section makes the environment reproducible from a fresh checkout.
- The final line points readers to the interactive API documentation.
- Save README.md.
- Confirm that the file contains the measured result table.
- Confirm that the file contains the Windows commands.
README formatting looks incomplete?
- Confirm that README.md is inside ai-retrieval-quality-gate.
- Check that the table retains every vertical separator.
- Check that the Windows commands remain inside the fenced code block.
Help me compare my README with the project target.
✔️ Awesome, I've got everything!
Your README now explains the measured result plus the local run process.
ⓧ I'd like to double check the full code
# AI Retrieval Quality Gate
A local retrieval API that compares a brittle exact-match baseline with TF-IDF ranking and blocks quality regressions with automated tests.
## Measured result
| Retriever | Hit@1 | MRR |
| --- | ---: | ---: |
| Exact-match baseline | 0.0 | 0.0 |
| TF-IDF ranking | 1.0 | 1.0 |
## Run on Windows
```text
py -V:3.14 -m venv .venv
.venv\Scripts\activate
py -m pip install -r requirements.txt
python evaluate.py
pytest
fastapi dev
```
Open `http://127.0.0.1:8000/docs` to test the API.
Publish and verify the gate
Git records the exact project state that passed locally. GitHub hosts that state so the workflow can test it independently.
- In GitHub, start a new repository.
- Enter ai-retrieval-quality-gate as the repository name.
- Set the repository visibility to public.
- Leave every repository initialization option disabled.
- Finish creating the empty repository.
GitHub should show instructions for pushing an existing local repository.
- Return to the activated terminal in the ai-retrieval-quality-gate folder.
- Initialize the repository and create the first commit by running these commands:
git init -b main
git add .
git commit -m "Build retrieval quality gate"
What do these commands create?
- The first command initializes Git on a branch named main.
- The second command stages the project files.
- The third command records the staged files as one commit.
When the commit succeeds, Git records your complete release gate on the main branch.
Git needs your identity?
Git may stop the commit when no author identity is configured. Replace the example values below with the name plus email you want attached to your commits.
- Replace Your Name with your Git author name.
- Replace you@example.com with your Git author email.
- Set the missing identity by running these commands:
git config --global user.name "Your Name"
git config --global user.email you@example.com
These settings attach your chosen identity to future commits.
- Retry the commit command from the previous code block.
Help me resolve my Git commit identity problem.
Git Credential Manager handles GitHub authentication for an HTTPS remote. The first authenticated push can open a browser sign-in flow.
- Copy the HTTPS repository URL from the empty GitHub repository page.
- Replace REMOTE-URL in the command below with the copied URL.
- Connect the repository and push the main branch by running these commands:
git remote add origin REMOTE-URL
git push -u origin main
What does the push do?
- The first command connects your local repository to the GitHub repository.
- The second command uploads the main branch.
- The upstream setting lets future pushes target the same branch.
- Complete the browser sign-in if Git Credential Manager prompts you.
- Return to the terminal after authentication completes.
You should see confirmation that the main branch was pushed. Your project is now public.
Push not reaching GitHub?
- Confirm that the remote URL belongs to the empty ai-retrieval-quality-gate repository.
- Complete any browser authentication request from Git Credential Manager.
- Check that your GitHub account can write to the repository.
Help me debug my GitHub push.
Before you open the workflow run, do you expect GitHub to reproduce the two passing local tests?
- In the GitHub repository, open the Actions tab.
- Open the newest Retrieval Evals run.
- Wait for the run to finish.
You should see a successful green run for Retrieval Evals. You have closed the loop. GitHub now blocks retrieval regressions with the same tests that passed locally.
Secret mission
Add an Abstention Threshold
Your retriever can rank relevant documents, but every ranking still needs enough evidence to deserve a result. Add an abstention threshold that returns an empty result list for unrelated queries, then protect that behavior with local and GitHub quality gates.
Clean Up Your Resources
Clean Up Your Resources
Both your local folder and public repository are free to keep under this project's setup. Decide whether to keep them available, pause your work, or delete them entirely.
Resources you used:
- Your local ai-retrieval-quality-gate folder with the source files and .venv environment.
- Your public GitHub repository with its GitHub Actions retrieval quality gate.
Keep everything running
No action needed. Choose this if you want to keep improving the retriever or demonstrate its quality gate.
- Keep the local ai-retrieval-quality-gate folder for future development.
- Keep the repository public so its standard hosted workflow continues to run at no cost.
- Continue adding labeled queries when you want to test new retrieval behavior.
Pause - I'll come back to this later
Shut down the local development process to free up memory. Your files and repository remain ready for your return.
- Close any Command Prompt windows that are using the project.
- Close the Visual Studio Code workspace from earlier.
- Leave the public repository unchanged until you are ready to continue.
What happens to the workflow?
The workflow only runs when a configured repository event occurs. Pausing your pushes means the quality gate stays idle.
Delete - I don't want to use this again
Remove the local folder and public repository when you no longer need the project. This clears both copies created during the walkthrough.
Deletion is permanent
Deleting both copies feels final because it removes your local work and remote project. A separate backup gives you a safe way back.
- Copy any project files you want to keep into a separate folder.
- Close the Visual Studio Code workspace from earlier.
- Switch back to the activated Command Prompt from earlier.
- Remove the local project folder by running these commands:
cd ..
rmdir /s /q ai-retrieval-quality-gate
What do these commands do?
- The first command moves Command Prompt to the folder that contains your project folder.
- The second command permanently removes ai-retrieval-quality-gate with its source files and .venv environment.
- Check the parent folder in File Explorer.
You should no longer see the ai-retrieval-quality-gate folder.
Still see the local folder?
- Close any remaining windows that are using files inside the project folder.
- Run the deletion commands again.
- Help me remove a locked Windows project folder.
- Open your ai-retrieval-quality-gate repository on GitHub.
- Select Settings.
- Scroll to the Danger Zone.
- Select Delete this repository.
- Read the deletion warnings.
- Confirm the warnings shown by GitHub.
- Enter ai-retrieval-quality-gate as the repository name.
- Complete the final deletion confirmation.
- Check your GitHub repository list.
You should no longer see the ai-retrieval-quality-gate repository or its workflow.
Nice Work!
Nice Work!
You did it! Your retrieval quality gate now ranks technical knowledge with measurable quality. Automated tests protect its ranking behavior.
What you learned:
- Compared an exact-match baseline with TF-IDF ranking powered by cosine similarity. This reproduced a brittle failure before improving retrieval for reordered queries.
- Built a labeled evaluation harness for repeatable quality checks. Used Hit@1 plus MRR to prove the ranked retriever reaches 1.0 on both metrics.
- Exposed ranked results through a documented FastAPI endpoint. Ran pytest in GitHub Actions to block retrieval regressions on every push.
- Secret Mission: Added an abstention threshold for unrelated queries. Protected that behavior with a third regression test.
Ready to quiz yourself?