Build a RAG API with FastAPI
Learn how to build a local AI pipeline that retrieves, augments, and generates answers from your own documents.
Introduction
⚡️ 30 Second Summary
Ever asked an AI a question about yourself and gotten a completely made-up answer? LLMs only know what they were trained on, not what's specific to you.
In this project, you'll build a RAG API using FastAPI, ChromaDB, and Ollama that answers personal questions by grounding AI responses in your own documents.
What You'll Build
You'll create a local REST API that implements the full Retrieval-Augmented Generation pipeline, from document storage to AI-generated answers, running entirely on your machine with zero cloud costs. RAG is used to build AI assistants that answer questions from internal documents, support tickets, or product catalogs without retraining models.
A question comes in through FastAPI, ChromaDB retrieves the most relevant context from your personal profile, and Ollama generates a grounded answer.
By the end of this project, you'll have:
- 📄 A RAG-assisted AI that you can ask questions about yourself and get accurate answers.
- 🔍 A ChromaDB vector database that retrieves relevant context using semantic search.
- 🚀 A FastAPI-powered /ask endpoint that implements the full RAG pipeline.
- 💎 Secret Mission: Extend your API into a multi-user AI directory where anyone can submit their profile and be queried.
Want a complete demo of how to do this project, from start to finish? Check out our 🎬 walkthrough with Pano
Are there any prerequisites?
There are no prerequisites, you can jump right in!
Not sure if this project is right for you? Check if it matches your goals
If you're up for a bit of a challenge, quiz yourself on the key concepts up ahead in this project.
This project is part of a series:
Set Up Your RAG Project
To build our RAG API, we need to get our development environment ready first. By the end of this project, you'll have a fully functional API that answers personal questions using AI, but first we need to install a few tools and dependencies.
In this step, you'll perform a manual RAG demo to understand the concept, then set up your Python project with everything it needs to build an AI-powered API.
In this step, get ready to:
- See RAG in action with a manual demo.
- Set up a Python project with a virtual environment.
- Install all project dependencies.
- Pull the nomic-embed-text embedding model.
Quick Start: Do you have Ollama installed?
This project uses Ollama to run AI models locally on your machine. Ollama is a free, open-source project that makes it easy to download and run large language models on your own hardware. No cloud subscriptions or API keys required.
✔️ I already have Ollama installed
Let's verify Ollama is running. Open your terminal and run the following command:
🍎 macOS/Linux
- Press Cmd + Space and type Terminal.
- Run the following command:
curl http://localhost:11434
🖼️ Windows
- Press the Win key and type PowerShell.
- Run the following command:
curl.exe http://localhost:11434
You should see Ollama is running in the response.
What is curl?
curl is a command-line tool for making web requests. Here, you're using it to send a request to Ollama's local server to check if it's responding.
✔️ I see Ollama is running
Great, Ollama is ready to go!
ⓧ I see a connection error
Ollama is installed but not running. You need to start it as a background service before any commands will work.
🍎 macOS
- Open Finder.
- Select Applications.
- Double-click Ollama.
You should see the Ollama icon appear in your menu bar at the top of the screen.
🖼️ Windows
- Press the Win key.
- Type Ollama.
- Press the Enter key.
You should see the Ollama icon appear in your system tray at the bottom-right of the screen.
🐧 Linux
- Open a new terminal window.
- Run ollama serve.
Keep this terminal open while you work through the project.
After starting Ollama, run the curl command again to verify you see Ollama is running.
Still stuck?
Get help with your error or share your error with the NextWork community!
ⓧ I need to install Ollama
No worries! Ollama is a free, open-source tool that runs AI models locally on your machine. No cloud API keys or subscriptions needed.
🍎 macOS
- Go to ollama.com and click Download.
- Click Download for macOS.
- Open the downloaded .dmg file and drag Ollama to your Applications folder.
- Open Ollama from your Applications folder.
Ollama will start running as a service in your menu bar. Close the chat window if it opens.
🖼️ Windows
- Go to ollama.com and click Download.
- Click Download for Windows.
- Open the downloaded .exe file and follow the installation prompts.
Ollama will start running as a service in your system tray. Close the chat window if it opens.
🐧 Linux
- Open your terminal app and run the following command:
curl -fsSL https://ollama.com/install.sh | sh
Now pull the AI model you'll use in this project:
🍎 macOS/Linux
- Press Cmd + Space and type Terminal.
- Run the following commands:
ollama pull qwen2.5:0.5b
🖼️ Windows
- Press the Win key and type PowerShell.
- Run the following commands:
ollama pull qwen2.5:0.5b
This downloads qwen2.5:0.5b, a small open-source chat model (~400MB) made by Alibaba. The "0.5b" means it has 500 million parameters, which is tiny compared to models like ChatGPT, but small enough to run on your laptop. You'll use it to generate answers from your personal data. Once the download finishes, verify Ollama is running:
🍎 macOS/Linux
curl http://localhost:11434
🖼️ Windows
curl.exe http://localhost:11434
You should see Ollama is running in the response. You're all set!
See RAG in Action
Before writing any code, let's demonstrate what RAG (Retrieval-Augmented Generation) actually does. You'll see the problem it solves and why it matters.
- Start a chat session with the qwen2.5:0.5b model:
ollama run qwen2.5:0.5b
- Ask the AI a personal question:
What are my career goals?
The AI has no idea who you are! It might make something up or give a generic answer, but it can't actually answer personal questions.
Now let's try again, but this time we'll add some personal context directly into your prompt.
- Fill out the variables below.
Based on this information about me:
"My name is [[YOURNAME="enter your name"]]. I'm learning cloud computing and AI.
My goal is to become a [[YOUR_CAREER_GOAL="e.g., DevOps engineer, cloud architect, AI engineer"]].
I enjoy [[YOUR_HOBBIES="e.g., hiking, playing guitar, reading sci-fi novels"]]."
What are my career goals?
- Copy and paste the above text into your chat.
- Send the message.
This time the AI gives you a grounded, accurate answer because you gave it context.
What just happened?
WOAH, you just performed RAG manually! You did three things:
- Retrieval - You found the relevant text (some personal info).
- Augmentation - You added that text to the prompt.
- Generation - The AI used that context to generate an accurate answer.
The rest of this project automates this exact process with code. How does RAG work in production systems?
- Type /bye to exit the chat session.
Set Up Your Python Project
Now let's create the project directory and set up a Python virtual environment to keep our dependencies isolated.
- In your terminal, create a new directory for the project and navigate into it:
🍎 macOS/Linux
mkdir ~/rag-api && cd ~/rag-api
🖼️ Windows
mkdir $HOME\rag-api; cd $HOME\rag-api
What does this command do?
This creates the rag-api folder in your home directory (~), then moves into it. The two commands are chained together so they run one after the other.
- Verify your Python version:
🍎 macOS/Linux
python3 --version
🖼️ Windows
python --version
You need Python 3.13 or higher for this project.
✔️ I see a version number
If you see Python 3.13 or higher, you're good to go.
ⓧ Command not found or version too low
If Python is not installed or the version is below 3.13:
🍎 macOS
- Install Python using Homebrew:
brew install python@3.13
- If you don't have Homebrew, install it first:
/bin/bash -c "$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)"
🖼️ Windows
- Download Python from python.org.
- Run the installer and make sure to check Add Python to PATH.
🐧 Linux
sudo apt update && sudo apt install python3.13 python3.13-venv
After installing, verify the version again with python3 --version.
Still stuck?
Get help with your error or share your error with the NextWork community!
- Create a virtual environment:
🍎 macOS/Linux
python3 -m venv venv
source venv/bin/activate
🖼️ Windows
python -m venv venv
venv\Scripts\activate
You should see (venv) appear at the beginning of your terminal prompt. This means you're now working inside the virtual environment.
Why use a virtual environment?
A virtual environment creates an isolated space for your project's dependencies. Without it, installing packages could conflict with other Python projects on your machine. The (venv) prefix in your terminal reminds you that packages are being installed in isolation. What problems do virtual environments solve?
Install Dependencies and Pull the Embedding Model
- Install all the Python packages we need:
pip install fastapi uvicorn chromadb ollama
What are these packages?
Here's what each package does:
- FastAPI - A modern Python web framework for building APIs. It auto-generates an interactive testing page called Swagger UI (you'll use this later).
- uvicorn - The server that runs your FastAPI application.
- ChromaDB - A vector database that stores and searches text using embeddings.
- ollama - The Python client for communicating with your local Ollama models.
What is a vector database?
- Pull the nomic-embed-text embedding model:
ollama pull nomic-embed-text
This downloads nomic-embed-text (~274MB), a model that converts text into numerical representations called embeddings. This is a different type of model to qwen2.5:0.5b. It doesn't chat, it converts text into numbers for search. The download may take a few minutes.
What is an embedding model?
An embedding model converts text into vectors that capture the meaning of the text.
Vectors are lists of numbers that represent text as coordinates in a high-dimensional space. Texts with similar meanings end up close together. This is the foundation of semantic search, where results are based on meaning rather than exact keyword matches.
Unlike chat models like qwen2.5 that generate responses, embedding models are purpose-built for search. They help you find the most relevant chunks of text for a given question. How do embeddings work?
- Verify the model is available:
ollama list
You should see both qwen2.5:0.5b and nomic-embed-text in the list. If you have other models installed from previous projects, that's fine. The important thing is that these two are present!
Don't see both models?
If nomic-embed-text isn't in the list, try running ollama pull nomic-embed-text again. If qwen2.5:0.5b is missing, run ollama pull qwen2.5:0.5b. Make sure Ollama is running before pulling models. My ollama pull command is failing
Your environment is ready with all dependencies installed and models downloaded. Next up, you'll create your personal profile document and build the knowledge base that powers the RAG pipeline.
Build Your Knowledge Base
You have all the tools installed and a working Python environment. The next piece of the puzzle is giving your AI something personal to work with. Right now, the AI only knows general facts from its training data. To answer questions about you, it needs your data.
In this step, you'll create a personal profile document, convert it into vector embeddings, and store it in a local database. This is the "Retrieval" part of RAG.
In this step, get ready to:
- Write a personal profile document.
- Build a Python script that loads, chunks, and stores your profile as embeddings.
- Run the script and verify your knowledge base is built.
Write Your Personal Profile
First, let's create the document that your AI will use to answer questions about you.
- In your terminal, create a new file called profile.txt:
🍎 macOS/Linux
touch profile.txt
🖼️ Windows
New-Item profile.txt
- Open profile.txt in your text editor by running this in your terminal:
🍎 macOS
open -a "TextEdit" profile.txt
🖼️ Windows
notepad profile.txt
This opens the file in Notepad.
🐧 Linux
xdg-open profile.txt
Or use your preferred editor like nano profile.txt or code profile.txt.
Pro tip: VS Code users
If you have VS Code installed, you can run code profile.txt in your terminal to open the file directly in VS Code.
- Add information about yourself. Here's a template to get you started:
My name is [[YOURNAME="enter your name"]].
I'm currently learning about cloud computing, AI, and DevOps.
My career goal is to become a [[YOUR_CAREER_GOAL="e.g., DevOps engineer, cloud architect, AI engineer"]].
I'm especially interested in [[YOUR_INTERESTS="e.g., automation, infrastructure as code, machine learning"]].
I have experience with [[YOUR_SKILLS="e.g., Python, JavaScript, AWS, Docker"]].
I'm currently building projects on NextWork to grow my hands-on skills.
For fun, I enjoy [[YOUR_HOBBIES="e.g., hiking, playing guitar, reading sci-fi novels"]].
A fun fact about me is [[YOUR_FUN_FACT="e.g., I once ran a marathon, I can solve a Rubik's cube in under a minute"]].
- Save the file and switch back to your terminal.
Why write a personal profile?
RAG systems work with all kinds of private documents: personal profiles, company knowledge bases, legal contracts, medical records, or product documentation. Any data that an AI wasn't trained on can be fed into a RAG pipeline to generate grounded answers.
When using cloud-based LLMs like ChatGPT or Claude, your data is sent to their servers. This is an important consideration when working with sensitive documents like medical records or internal company data.
Running LLMs locally with Ollama means your data never leaves your machine, which allows you to safely process personal or confidential information.
Create the Knowledge Base Script
Now you'll write a Python script that takes your profile document, splits it into chunks, generates embeddings for each chunk, and stores them in ChromaDB.
- In your terminal, create a new file called build_knowledge_base.py:
🍎 macOS/Linux
touch build_knowledge_base.py
🖼️ Windows
New-Item build_knowledge_base.py
- Run ls to confirm the file was created:
ls
- Open build_knowledge_base.py in your text editor by running this in your terminal:
🍎 macOS
open -a "TextEdit" build_knowledge_base.py
🖼️ Windows
notepad build_knowledge_base.py
🐧 Linux
xdg-open build_knowledge_base.py
- Paste the following code into your text editor.
import chromadb
from chromadb.utils.embedding_functions.ollama_embedding_function import (
OllamaEmbeddingFunction,
)
# Load the profile document
with open("profile.txt", "r") as f:
text = f.read()
# Split into chunks by paragraph - each blank line becomes a split point
# strip() removes extra whitespace, and the if-check skips empty chunks
chunks = [chunk.strip() for chunk in text.split("\n\n") if chunk.strip()]
print(f"Loaded {len(chunks)} chunks from profile.txt")
What does this code do?
This loads your profile.txt file and splits it into chunks by paragraph. Smaller chunks help the AI find the most relevant section rather than dumping the entire document into the prompt. Why do we split documents into chunks?
- Next, paste the following code below what you just added.
# Initialize ChromaDB - PersistentClient saves data to disk so it survives restarts
client = chromadb.PersistentClient(path="./chroma_db")
# Connect to Ollama's embedding model to convert text into vectors
ef = OllamaEmbeddingFunction(
model_name="nomic-embed-text",
url="http://localhost:11434", # Ollama's default local address
)
# Create (or reuse) a collection - like a table in a database
collection = client.get_or_create_collection(
name="personal_profile",
embedding_function=ef, # Tells ChromaDB how to convert text to vectors
)
What does this code do?
This creates a ChromaDB database that saves to disk, connects it to nomic-embed-text for generating embeddings, and creates a collection called personal_profile. A collection is like a table in a relational database.
Here's how the pieces fit together: the OllamaEmbeddingFunction tells ChromaDB to call your local Ollama server whenever it needs to convert text into vectors. ChromaDB handles the storage and search, but Ollama's nomic-embed-text model does the actual embedding work behind the scenes.
- Finally, paste the following code below what you just added.
# Add chunks to the collection - ChromaDB automatically generates embeddings
collection.add(
ids=[f"chunk{i}" for i in range(len(chunks))], # Unique ID for each chunk
documents=chunks, # The actual text content
metadatas=[{"source": "profile", "chunk_index": i} for i in range(len(chunks))],
)
print(f"Added {len(chunks)} chunks to the 'personal_profile' collection.")
print("Knowledge base built successfully!")
What does this code do?
When you call collection.add(), ChromaDB automatically sends each chunk to nomic-embed-text, converts it into a vector, and stores both the text and vector together.
✔️ I built it step by step
Your file should now look like the screenshot below. If something doesn't look right, check the Full code reference tab.
📋 Full code reference
If you'd prefer to paste the complete file:
import chromadb
from chromadb.utils.embedding_functions.ollama_embedding_function import (
OllamaEmbeddingFunction,
)
# Load the profile document
with open("profile.txt", "r") as f:
text = f.read()
# Split into chunks by paragraph - each blank line becomes a split point
# strip() removes extra whitespace, and the if-check skips empty chunks
chunks = [chunk.strip() for chunk in text.split("\n\n") if chunk.strip()]
print(f"Loaded {len(chunks)} chunks from profile.txt")
# Initialize ChromaDB - PersistentClient saves data to disk so it survives restarts
client = chromadb.PersistentClient(path="./chroma_db")
# Connect to Ollama's embedding model to convert text into vectors
ef = OllamaEmbeddingFunction(
model_name="nomic-embed-text",
url="http://localhost:11434", # Ollama's default local address
)
# Create (or reuse) a collection - like a table in a database
collection = client.get_or_create_collection(
name="personal_profile",
embedding_function=ef, # Tells ChromaDB how to convert text to vectors
)
# Add chunks to the collection - ChromaDB automatically generates embeddings
collection.add(
ids=[f"chunk{i}" for i in range(len(chunks))], # Unique ID for each chunk
documents=chunks, # The actual text content
metadatas=[{"source": "profile", "chunk_index": i} for i in range(len(chunks))],
)
print(f"Added {len(chunks)} chunks to the 'personal_profile' collection.")
print("Knowledge base built successfully!")
- Save the file.
- Switch back to your terminal and run the script:
python build_knowledge_base.py
✔️ Knowledge base built successfully
You should see output confirming how many chunks were loaded and stored.
ⓧ I see an error
That's okay! Let's troubleshoot:
- Check that Ollama is running. A connection error means the embedding model can't be reached. How do I check if Ollama is running?
- Check that profile.txt is in the same directory as build_knowledge_base.py. A FileNotFoundError means the script can't find your file.
- If you see a duplicate ID error, delete the ./chroma_db folder and run the script again.
Still stuck?
Get help with your error or share your error with the NextWork community!
What is happening behind the scenes?
When you run this script, ChromaDB sends each text chunk to nomic-embed-text, which converts it into a 768-dimensional vector. That means each chunk of text is represented as a list of 768 numbers. Each number captures a different aspect of the text's meaning, like topic, tone, or context.
Embedding models range from 384 dimensions (smaller, faster) to 1536 dimensions (larger, more precise). At 768, nomic-embed-text balances search quality with speed.
💡 How does the search actually work?
These vectors are stored locally at ./chroma_db. When someone asks a question, that question is also converted into a vector, and ChromaDB finds the chunks whose vectors are closest in that high-dimensional space. This is semantic search in action! Instead of matching exact keywords, it finds content with the closest meaning. How does semantic search work?
Your knowledge base is built and searchable. Next, you'll create the API that ties everything together, from question to retrieval to AI-generated answer.
Build Your RAG API
You've got a vector database loaded with your personal profile, and you've seen that it can retrieve relevant chunks. Now it's time to build the API that brings the full RAG pipeline together. When someone sends a question to your API, it will automatically retrieve context, augment the prompt, and generate a grounded answer.
This is the final piece. After this step, you'll have a working API that anyone can call to ask questions about you.
What is an API?
An API (Application Programming Interface) is a way for programs to talk to each other. Instead of clicking buttons in a user interface, code sends a request to a specific URL and gets structured data back. For example, when a weather app shows today's forecast, it's calling a weather API behind the scenes. In this step, you'll build an API that accepts a question and returns an AI-generated answer.
In this step, get ready to:
- Create a FastAPI application with a /ask endpoint.
- Test your API using the built-in Swagger UI.
Create Your FastAPI Application
Now you'll write the main API file that implements the complete RAG pipeline in a single endpoint.
- In your terminal, create a new file called main.py:
🍎 macOS/Linux
touch main.py
🖼️ Windows
New-Item main.py
- Open main.py in your text editor by running this in your terminal:
🍎 macOS
open -a "TextEdit" main.py
🖼️ Windows
notepad main.py
🐧 Linux
xdg-open main.py
- Paste the following code into your text editor.
from fastapi import FastAPI
import ollama
import chromadb
from chromadb.utils.embedding_functions.ollama_embedding_function import (
OllamaEmbeddingFunction,
)
app = FastAPI() # Create the FastAPI application
# Connect to the same ChromaDB collection you built in Step 2
client = chromadb.PersistentClient(path="./chroma_db")
ef = OllamaEmbeddingFunction(
model_name="nomic-embed-text",
url="http://localhost:11434",
)
collection = client.get_or_create_collection(
name="personal_profile",
embedding_function=ef,
)
What does this code do?
This creates a FastAPI application and connects to the same ChromaDB collection you built in Step 2. When the server starts, it will have immediate access to your knowledge base. Why does FastAPI use a GET endpoint here?
- Next, paste the following code below what you just added.
@app.get("/ask") # This creates a GET endpoint at /ask
def ask(question: str): # FastAPI automatically reads "question" from the URL query string
# Step 1: RETRIEVE - search ChromaDB for the 2 most relevant chunks
results = collection.query(
query_texts=[question], # ChromaDB converts this to a vector and finds similar chunks
n_results=2, # Return the top 2 matches
)
# Combine the matching chunks into a single string
context = "\n\n".join(results["documents"][0])
# Step 2: AUGMENT - build a prompt that includes the retrieved context
augmented_prompt = f"""Use the following context to answer the question.
If the context doesn't contain relevant information, say so.
Context:
{context}
Question: {question}"""
# Step 3: GENERATE - send the augmented prompt to the local LLM
response = ollama.chat(
model="qwen2.5:0.5b",
messages=[{"role": "user", "content": augmented_prompt}],
)
# Return the answer along with the context so users can verify the source
return {
"question": question,
"answer": response["message"]["content"],
"context_used": results["documents"][0],
}
What does this code do?
This single endpoint implements the three steps of RAG that you performed manually in Step 1. When someone sends a question, it retrieves the 2 most relevant chunks from ChromaDB, augments the prompt by combining those chunks with the question, and generates a grounded answer using qwen2.5:0.5b. The response includes the context that was used so you can verify the AI's sources.
✔️ I built it step by step
Your file should now look like the screenshot below. If something doesn't look right, check the Full code reference tab.
📋 Full code reference
If you'd prefer to paste the complete file:
from fastapi import FastAPI
import ollama
import chromadb
from chromadb.utils.embedding_functions.ollama_embedding_function import (
OllamaEmbeddingFunction,
)
app = FastAPI() # Create the FastAPI application
# Connect to the same ChromaDB collection you built in Step 2
client = chromadb.PersistentClient(path="./chroma_db")
ef = OllamaEmbeddingFunction(
model_name="nomic-embed-text",
url="http://localhost:11434",
)
collection = client.get_or_create_collection(
name="personal_profile",
embedding_function=ef,
)
@app.get("/ask") # This creates a GET endpoint at /ask
def ask(question: str): # FastAPI automatically reads "question" from the URL query string
# Step 1: RETRIEVE - search ChromaDB for the 2 most relevant chunks
results = collection.query(
query_texts=[question], # ChromaDB converts this to a vector and finds similar chunks
n_results=2, # Return the top 2 matches
)
# Combine the matching chunks into a single string
context = "\n\n".join(results["documents"][0])
# Step 2: AUGMENT - build a prompt that includes the retrieved context
augmented_prompt = f"""Use the following context to answer the question.
If the context doesn't contain relevant information, say so.
Context:
{context}
Question: {question}"""
# Step 3: GENERATE - send the augmented prompt to the local LLM
response = ollama.chat(
model="qwen2.5:0.5b",
messages=[{"role": "user", "content": augmented_prompt}],
)
# Return the answer along with the context so users can verify the source
return {
"question": question,
"answer": response["message"]["content"],
"context_used": results["documents"][0],
}
- Save the file and switch back to your terminal.
Test with Swagger UI
FastAPI automatically generates interactive documentation for your API. Let's start the server and try it out.
- In your terminal, start the API server:
uvicorn main:app --reload
You should see output indicating the server is running at http://127.0.0.1:8000.
What does --reload do?
The --reload flag tells uvicorn to automatically restart the server whenever you change your code. This is useful during development because you don't have to manually stop and restart the server every time you make an edit.
- Open your browser and go to http://127.0.0.1:8000/docs.
This is FastAPI's built-in Swagger UI, an interactive API playground that auto-generates from your code. You can test your endpoints directly in the browser without any additional tools.
- Click on the GET /ask endpoint to expand it.
- Click Try it out.
- In the question field, enter a personal question like What is my name?.
- Click Execute.
✔️ I see a JSON response
You just got your RAG API to answer a personal question! Take a look at the JSON response:
- answer contains the AI's response, grounded in your actual profile data.
- context_used shows the exact chunks that ChromaDB retrieved from your knowledge base. This is the context that was injected into the prompt before the AI generated its answer.
Remember when the AI couldn't answer personal questions back in Step 1? Now it can, because your API is automatically retrieving the right context and augmenting the prompt before the AI generates a response. That's the full RAG pipeline running end-to-end!
ⓧ I see an error or wrong answer
That's okay! Let's troubleshoot:
- If the response says the context doesn't contain relevant information, your question might not match your profile content closely enough. Try asking something that directly relates to what you wrote in profile.txt.
- If you get a connection error, make sure Ollama is running. How do I check if Ollama is running?
- If the server isn't responding, check your first terminal window to make sure uvicorn is still running.
Still stuck?
Get help with your error or share your error with the NextWork community!
Why is Swagger UI useful?
Swagger UI is one of the biggest advantages of FastAPI over other frameworks like Flask. It auto-generates interactive documentation from your code's type hints, so anyone can explore and test your API without reading any documentation. This is a major win for both development and collaboration.
- Stop the server by pressing Ctrl+C in your terminal.
Your RAG API is live and working. You've built a complete pipeline that retrieves personal context, augments the prompt, and generates grounded answers. Ready for a challenge? The Secret Mission takes your API to the next level!
Secret mission
You've built a RAG API that answers questions about you, but what if other people could add their profiles too? In this secret mission, you'll extend the API so that multiple users can submit their own personal profiles, turning your RAG API into a multi-user personal AI directory.
In this secret mission, get ready to:
- Add a POST /documents endpoint for submitting new profiles.
- Update the /ask endpoint to filter by user.
- Test the multi-user directory with a second profile.
Build a Multi-User AI Directory
Clean Up Your Resources
Clean Up Your Resources
Now that you've built and tested your RAG API, let's decide what to do with the resources on your machine.
Your RAG API project is great portfolio material.
Consider keeping it to showcase your AI and API development skills to employers.
Resources to manage:
- Uvicorn server (currently running).
- Python virtual environment.
- ChromaDB data directory (./chroma_db).
- Project files (rag-api directory).
🟢 Keep everything
No action needed! Your project files and database are stored locally with no ongoing costs. Keep the project for Part 3 of the series where you'll containerize this API with Docker.
🟡 Stop the server but keep files
If you want to stop the running services but keep your project files for later:
- Stop the uvicorn server by pressing Ctrl+C in the terminal where it's running.
- Deactivate the virtual environment:
deactivate
To restart later, navigate to your rag-api directory, activate the virtual environment, and run uvicorn main:app --reload again.
🔴 Delete everything
If you want to remove the entire project:
- Stop the uvicorn server by pressing Ctrl+C.
- Deactivate the virtual environment:
deactivate
- Delete the project directory:
🍎 macOS/Linux
cd .. && rm -rf rag-api
🖼️ Windows
cd .. ; Remove-Item -Recurse -Force rag-api
- Optionally remove the nomic-embed-text model to free up 274MB of disk space:
ollama rm nomic-embed-text
Should I remove the models?
If you're continuing with the AI Fundamentals series, you'll need both qwen2.5:0.5b and nomic-embed-text for upcoming projects. We recommend keeping them installed.
That's a wrap!
That's a wrap!
Nice work! 🚀 You've built a fully functional RAG API that answers personal questions using AI, running entirely on your local machine.
You've learned how to:
- 🔍 Perform RAG (Retrieval-Augmented Generation) both manually and with code.
- 📄 Create a personal knowledge base using ChromaDB and vector embeddings.
- 🚀 Build a REST API with FastAPI that implements a full RAG pipeline.
- 🧠 Use nomic-embed-text for semantic search and qwen2.5:0.5b for AI response generation.
- 💎 Extend the API into a multi-user AI directory with dynamic document ingestion.
Congratulations! 🎉
You now have a working RAG API that retrieves relevant context from a vector database, augments the prompt, and generates grounded answers. This is the same architecture behind many AI products today.
Ready to quiz yourself? 💪
p.s. Does it say "Still tasks to complete!" at the bottom of the screen?
This means you still have screenshots left to upload, or questions left to answer!
- Press Ctrl+F (Windows) or Command+F (Mac) on your keyboard.
- Search for the text Return to later.
- Jump straight to your incomplete tasks!
- 🙋♀️ Still stuck? Ask the community!