Build a Self-Updating RAG Assistant
Sync live GitHub Actions docs into a cited, source-aware RAG assistant.
Introduction
30 Second Summary
When you rely on product documentation at work, yesterday's correct answer can become today's stale advice. The problem is easy to miss when an assistant keeps searching an old copy.
In this project, you will build a self-updating retrieval-augmented generation assistant that answers questions from current GitHub Actions documentation. It will refresh changed sources while keeping official documentation distinct from client guidance.
What You'll Build
You'll ask where workflow files belong and receive a current answer with official evidence displayed beside it.
By the end of this project, you'll have:
- A live source probe that displays three current document titles, repository paths, and 12-character Git blob SHA prefixes from GitHub's official documentation.
- An incremental sync report backed by a persistent local Qdrant database, showing which sources were updated, skipped, removed, deleted, or upserted.
- A source-aware assistant in Streamlit that answers from retrieved sections and labels each result as official documentation or a client assumption.
- Secret Mission: Run a controlled source retirement and recovery drill that proves removed vectors disappear and restored sources return.
Are there any prerequisites?
You'll need an internet connection plus a Mac with Python 3.10 through 3.14, Visual Studio Code, Git, and GitHub access. You'll also need permission to create an OpenAI API key with paid usage protected by a $10 ceiling.
Before We Start
Before the hands-on work begins, take a moment to lock in the client scenario behind this project. Your local RAG assistant must keep official GitHub Actions documentation separate from private client guidance so every answer presents its sources honestly.
Set Up the Live Docs RAG Workspace
Your live RAG assistant depends on changing GitHub Actions documentation. A reproducible workspace keeps that foundation consistent whenever you return to the project.
In this step, you will verify Python compatibility. You will prepare local Qdrant storage plus protected access to the OpenAI API before confirming that Streamlit runs.
In this step, get ready to:
- Verify Python compatibility plus the macOS developer tools.
- Create an isolated workspace with pinned dependencies plus project files.
- Protect the API budget before launching Streamlit.
Verify Python and developer tools
The pinned packages support Python 3.10 through 3.14. The Command Line Tools for Xcode provide development utilities used by the local environment.
- Press Cmd+Space to open Spotlight.
- Type Terminal into Spotlight.
- Press Enter to open Terminal.
- Check the installed Python version by running this command:
python3 --version
What does this command check?
The command prints the version connected to python3. This project needs a version from 3.10 through 3.14.
✔️ I see a supported version
Your Python interpreter supports every pinned package used in this project.
ⓧ I see an older version
Your current interpreter is below Python 3.10. Install a supported release before creating the virtual environment.
- Open the official Python releases for macOS page.
- Download a stable Python 3.14 macOS installer.
- Complete the macOS installer.
- Close the current Terminal window.
- Press Cmd+Space to open Spotlight.
- Type Terminal into Spotlight.
- Press Enter to start a new shell session.
- Check the active Python version by running this command:
python3 --version
What should the new check prove?
The new shell session reloads the available Python commands. Continue when the printed version is from 3.10 through 3.14.
ⓧ Command not found
Your shell cannot find Python 3. The official macOS installer adds a compatible interpreter.
- Open the official Python releases for macOS page.
- Download a stable Python 3.14 macOS installer.
- Complete the macOS installer.
- Press Cmd+Space to open Spotlight.
- Type Terminal into Spotlight.
- Press Enter to start a new shell session.
- Confirm that Python is available by running this command:
python3 --version
What should the installation change?
Your shell should now recognize python3. Continue when the printed version is inside the supported range.
- Check the active developer tools path by running this command:
xcode-select --print-path
What does this command check?
The command asks macOS for the active developer tools directory. A printed path confirms that the tools are available.
✔️ I see a developer tools path
Your Mac can reach the development utilities needed by the project environment.
ⓧ Developer tools are unavailable
macOS provides an installation prompt for the missing tools. Complete that flow before preparing the workspace.
- Start the Command Line Tools installation by running this command:
xcode-select --install
What does this command start?
macOS opens the Command Line Tools installation flow. The installer adds the development utilities without changing your project files.
- Complete the macOS installation prompt.
- Confirm the installed developer tools path by running this command:
xcode-select --print-path
What should the second check show?
The command should now print the active developer tools directory. That path confirms the installation succeeded.
Create the isolated workspace
A virtual environment keeps this project's packages separate from system-wide packages. Visual Studio Code gives the files plus terminal one shared location.
- Press Cmd+Space to open Spotlight.
- Type Visual Studio Code into Spotlight.
- Press Enter to open Visual Studio Code.
- Open the folder picker from the top menu.
- Select your Desktop as the location.
- Create a folder named Self-Updating GitHub Actions Docs RAG.
You now have a dedicated Desktop folder for every project file plus generated local resource.
- Select the Self-Updating GitHub Actions Docs RAG folder as the workspace.
- Confirm that the folder name appears at the top of the Explorer sidebar.
Wrong folder in Explorer?
Use the folder picker again if Explorer shows your whole Desktop. Select only the Self-Updating GitHub Actions Docs RAG folder.
Use Help me select the correct Visual Studio Code workspace..
The first file pins each library version. Exact pins keep future installations aligned with the code used in this project.
- Use the Explorer sidebar's new-file control to create requirements.txt.
- Paste the following dependency pins into requirements.txt:
openai==3.26.1
qdrant-client==1.19.1
streamlit==1.65.0
What does this file control?
- The openai==3.26.1 pin provides the OpenAI Python SDK.
- The qdrant-client==1.19.1 pin provides the persistent local vector client.
- The streamlit==1.65.0 pin provides the browser interface.
- Save requirements.txt.
- Confirm that requirements.txt appears in the Explorer sidebar.
Requirements file missing?
Confirm that Explorer shows the project folder. Recreate requirements.txt inside that folder if it was saved elsewhere.
Use Help me locate requirements.txt in my workspace..
Generated environments plus vector data should stay outside future commits. The ignore file identifies each local artifact that can be recreated.
- Use the Explorer sidebar's new-file control to create .gitignore.
- Paste the following exclusions into .gitignore:
.venv/
__pycache__/
qdrant_data/
manifest.json
What does this file protect?
The file excludes the virtual environment plus Python cache files. It also excludes the generated Qdrant database plus source manifest.
- Save .gitignore.
- Confirm that .gitignore appears in the Explorer sidebar.
Gitignore name looks wrong?
Confirm that the filename begins with a full stop. Remove any extra extension after .gitignore.
Use Help me correct the .gitignore filename..
The client overlay keeps private guidance separate from official GitHub documentation. Its headings also provide the controlled source you will update later.
- Use the Explorer sidebar's new-file control to create assumptions.md.
- Paste the following client guidance into assumptions.md:
# Client Assumptions
## Freshness target
This portfolio demo assumes the client wants the private guidance and vendor documentation checked every 15 minutes by an external scheduler.
## Source policy
Official GitHub documentation is authoritative for GitHub Actions behavior. This client overlay may add internal recommendations, but the assistant must label them as assumptions and must not present them as GitHub policy.
## Unsupported questions
When retrieved evidence does not support an answer, the assistant should say that the synchronized sources do not contain enough information.
What does this overlay represent?
The freshness target models a private client requirement. The source policy prevents that requirement from being presented as official GitHub policy.
The unsupported-question rule gives the future assistant a clear boundary when evidence is missing.
- Save assumptions.md.
- Confirm that assumptions.md appears in the Explorer sidebar.
Overlay formatting looks different?
Compare the heading markers plus blank lines with the reference below. Keep every policy sentence unchanged.
Use Help me compare assumptions.md with the required structure..
✔️ Awesome, I've got everything!
Your three starter files are saved in the project workspace.
ⓧ I'd like to double check the full code
- Compare each saved file with its complete reference below.
openai==3.26.1
qdrant-client==1.19.1
streamlit==1.65.0
Requirements checkpoint
The file contains exactly three dependency pins. Each line matches the package version used by the project.
.venv/
__pycache__/
qdrant_data/
manifest.json
Gitignore checkpoint
The file contains exactly four local exclusions. Each entry matches a generated project artifact.
# Client Assumptions
## Freshness target
This portfolio demo assumes the client wants the private guidance and vendor documentation checked every 15 minutes by an external scheduler.
## Source policy
Official GitHub documentation is authoritative for GitHub Actions behavior. This client overlay may add internal recommendations, but the assistant must label them as assumptions and must not present them as GitHub policy.
## Unsupported questions
When retrieved evidence does not support an answer, the assistant should say that the synchronized sources do not contain enough information.
Assumptions checkpoint
The overlay contains the freshness target plus source policy. It also contains the unsupported-question rule.
- Open the integrated terminal from the Visual Studio Code top menu.
- Confirm that the terminal starts inside the Self-Updating GitHub Actions Docs RAG workspace.
The terminal now runs commands inside the folder containing your three starter files.
- Create the isolated Python environment by running this command:
python3 -m venv .venv
What does this command create?
Python creates an environment inside .venv. That folder receives its own interpreter plus package installation area.
- Confirm that .venv appears in the Explorer sidebar.
Virtual environment missing?
Confirm that the terminal is inside the project workspace. Check that the Python version command succeeded earlier.
Use Help me create the .venv environment..
- Activate the new virtual environment by running this command:
source .venv/bin/activate
What does activation change?
Activation points package commands at .venv for the current terminal session. The prompt should now begin with (.venv).
- Confirm that (.venv) appears at the start of the terminal prompt.
Environment not activated?
Confirm that .venv exists in the Explorer sidebar. Run the activation command from the project folder.
Use Help me activate the .venv environment..
- Install the pinned project dependencies by running this command:
python3 -m pip install -r requirements.txt
What does this command install?
The active environment reads each exact pin from requirements.txt. It installs the OpenAI SDK plus the Qdrant client plus Streamlit.
- Wait for the terminal to return to the (.venv) prompt.
- Confirm that the installation finishes without a package error.
Dependency installation failed?
Confirm that the prompt begins with (.venv). Compare every line in requirements.txt with the full-code reference.
Use Help me troubleshoot the pinned dependency installation..
Protect API access and launch Streamlit
The later synchronization plus answer flow makes token-metered OpenAI requests. An API key provides access while the monthly hard limit places a spending ceiling around the project.
Protect the key
Creating a credential deserves a careful pause. Store the key securely because its complete value is available only during creation.
Keep the key outside every project file. The shell environment makes it available without placing the secret in code.
- Sign in to OpenAI Platform in your browser.
- Go to the API keys area.
- Create an API key for this workspace.
- Store the displayed key in a password manager.
- Close the creation dialog after the key is stored.
Your key is now stored away from the project. The next control limits the amount that API requests can charge.
Set the limit before testing
Billing controls can feel high stakes. You are setting a $10 monthly ceiling before this project makes a paid request.
Hard-limit enforcement is not instantaneous. Recorded spend can slightly exceed the configured amount.
- Open Organization limits in OpenAI Platform.
- Select Spend.
- Select Edit spend limit.
- Enter 10 in Monthly spend limit.
- Turn on Enforce a hard limit.
- Select Save.
The saved control protects future embedding plus generation tests with the configured monthly ceiling.
- Return to the activated Visual Studio Code terminal.
- Replace your_api_key_here with the stored secret in the command below.
- Export the completed key command in the active shell:
export OPENAI_API_KEY="your_api_key_here"
What does this command change?
The shell stores the credential under OPENAI_API_KEY for programs launched from this terminal. The command does not write the secret into a project file.
A successful export returns to the prompt without printing the key. Keep the terminal line containing the secret out of screenshots.
Concerned about key exposure?
Confirm that the key appears only in the active shell command. Remove it immediately from any project file if it was pasted there.
Use Help me verify that OPENAI_API_KEY is stored safely..
Before you launch it, consider whether Streamlit will use the system interpreter or the active virtual environment.
- Launch the Streamlit Hello app from the activated environment by running this command:
python -m streamlit hello
What does this command prove?
The active Python interpreter loads Streamlit from .venv. Streamlit starts its local server before opening the Hello app in your browser.
- Confirm that the Streamlit Hello app opens in a browser tab.
- Keep the Streamlit process running for the screenshot checkpoint.
That is the environment check complete. Your isolated Streamlit installation now runs successfully from the project workspace.
Hello app did not open?
Confirm that the terminal showed (.venv) before the launch command. Open the local address printed in the terminal if the browser did not open automatically.
Use Help me launch the Streamlit Hello app..
Your protected workspace is ready for live documentation. Next, you will inspect current GitHub document versions before building retrieval.
Inspect Live GitHub Docs Versions
Your local workspace is ready for the project code. The assistant still has no way to prove which documentation version it reads.
Each official file exposes a Git blob SHA through the GitHub Contents API. In this step, you will fetch three live GitHub Actions documents. You will print their current titles, repository paths, and version identifiers.
In this step, get ready to:
- Define a catalog for three official GitHub Actions documentation files.
- Decode each file's Base64 content and retain its version metadata.
- Prove that the documentation source is live with a terminal report.
Define the official source catalog
A stable source catalog gives each remote document one repository path. The path lets every later sync compare the same source over time.
- Switch back to the Visual Studio Code window from the previous step.
- In the Explorer sidebar, select the folder that contains requirements.txt.
- Click the new-file control beside that folder.
- Enter rag_core.py as the file name.
- Press Enter to create the file.
- Build the source catalog in rag_core.py by pasting this code:
from base64 import b64decode
import json
from urllib.parse import quote
from urllib.request import Request, urlopen
GITHUB_API_ROOT = "https://api.github.com/repos/github/docs/contents"
GITHUB_API_VERSION = "2026-03-10"
REMOTE_SOURCES = [
{
"path": "content/actions/reference/workflows-and-actions/workflow-syntax.md",
"label": "Workflow syntax for GitHub Actions",
},
{
"path": "content/actions/reference/workflows-and-actions/workflow-commands.md",
"label": "Workflow commands for GitHub Actions",
},
{
"path": "content/actions/reference/workflows-and-actions/dependency-caching.md",
"label": "Dependency caching reference",
},
]
def extract_title(text, fallback):
for line in text.splitlines()[:20]:
if line.startswith("title:"):
return line.split(":", 1)[1].strip().strip('"')
return fallback
What does this code establish?
- The imports provide the standard-library tools needed to request remote files.
- The JSON helper reads the API response into a Python object.
- The GITHUB_API_ROOT constant points to the public repository contents endpoint.
- The GITHUB_API_VERSION constant selects API version 2026-03-10.
- The REMOTE_SOURCES list records the three documentation paths that this project tracks.
- The extract_title() helper searches the first 20 lines for a frontmatter title.
- Save rag_core.py by pressing Cmd+S on macOS or Ctrl+S on Windows.
- Confirm rag_core.py now shows three entries inside REMOTE_SOURCES.
Seeing incomplete code highlighting?
- Check that every opening square bracket has a matching closing square bracket.
- Check that each source dictionary ends with a closing curly brace.
- Compare each repository path with the code block because one missing character changes the requested file.
- Help me check my source catalog for a Python syntax mistake.
Fetch and normalize each document
The API returns file contents as encoded text alongside repository metadata. A normalized source record keeps that live text together with its path, SHA, URL, type, and title.
- In rag_core.py, place your cursor below the final return fallback line.
- Add the GitHub source fetcher by pasting this code:
def fetch_github_source(source):
encoded_path = quote(source["path"], safe="/")
url = f"{GITHUB_API_ROOT}/{encoded_path}?ref=main"
request = Request(
url,
headers={
"Accept": "application/vnd.github+json",
"X-GitHub-Api-Version": GITHUB_API_VERSION,
"User-Agent": "live-docs-rag-demo",
},
)
with urlopen(request, timeout=30) as response:
payload = json.load(response)
text = b64decode(payload["content"]).decode("utf-8")
return {
"source_path": payload["path"],
"source_version": payload["sha"],
"source_url": payload["html_url"],
"source_kind": "official",
"title": extract_title(text, source["label"]),
"text": text,
}
How does the fetcher work?
- The encoded path becomes a request URL against the repository's main branch.
- The Accept header requests GitHub's JSON representation.
- The X-GitHub-Api-Version header keeps the request on the selected API version.
- The User-Agent header identifies this learning application to GitHub.
- The decoder converts the response's content field into readable Markdown.
- The returned dictionary keeps the blob sha as source_version.
- Save rag_core.py.
- Confirm the fetch_github_source() function ends with a dictionary containing six normalized source fields.
Does the function look misaligned?
- Keep the request headers indented inside the headers dictionary.
- Keep payload = json.load(response) indented inside the response context.
- Keep the decoded text line outside that response context.
- Help me fix the indentation in fetch_github_source().
✔️ Awesome, I've got everything!
Great. Double-check that you saved rag_core.py.
ⓧ I'd like to double check the full code
from base64 import b64decode
import json
from urllib.parse import quote
from urllib.request import Request, urlopen
GITHUB_API_ROOT = "https://api.github.com/repos/github/docs/contents"
GITHUB_API_VERSION = "2026-03-10"
REMOTE_SOURCES = [
{
"path": "content/actions/reference/workflows-and-actions/workflow-syntax.md",
"label": "Workflow syntax for GitHub Actions",
},
{
"path": "content/actions/reference/workflows-and-actions/workflow-commands.md",
"label": "Workflow commands for GitHub Actions",
},
{
"path": "content/actions/reference/workflows-and-actions/dependency-caching.md",
"label": "Dependency caching reference",
},
]
def extract_title(text, fallback):
for line in text.splitlines()[:20]:
if line.startswith("title:"):
return line.split(":", 1)[1].strip().strip('"')
return fallback
def fetch_github_source(source):
encoded_path = quote(source["path"], safe="/")
url = f"{GITHUB_API_ROOT}/{encoded_path}?ref=main"
request = Request(
url,
headers={
"Accept": "application/vnd.github+json",
"X-GitHub-Api-Version": GITHUB_API_VERSION,
"User-Agent": "live-docs-rag-demo",
},
)
with urlopen(request, timeout=30) as response:
payload = json.load(response)
text = b64decode(payload["content"]).decode("utf-8")
return {
"source_path": payload["path"],
"source_version": payload["sha"],
"source_url": payload["html_url"],
"source_kind": "official",
"title": extract_title(text, source["label"]),
"text": text,
}
Run the live source probe
A probe gives you a small visible test of the ingestion path. It calls the fetcher once for every configured source.
- In the Explorer sidebar, select the folder that contains rag_core.py.
- Click the new-file control beside that folder.
- Enter source_probe.py as the file name.
- Press Enter to create the file.
- Build the probe in source_probe.py by pasting this code:
from rag_core import REMOTE_SOURCES, fetch_github_source
for configured_source in REMOTE_SOURCES:
source = fetch_github_source(configured_source)
print(
f"{source['title']} | {source['source_version'][:12]} | "
f"{source['source_path']}"
)
What does the probe report?
- The import reuses the source catalog and fetcher from rag_core.py.
- The loop requests every configured document once.
- The first printed value is the extracted document title.
- The second printed value is the first 12 characters of the current blob SHA.
- The final printed value is the document's repository path.
- Save source_probe.py.
- Return to the activated terminal from the previous step.
- If the Streamlit Hello app is still running, press Ctrl+C to stop it.
Before you run the probe, how many source records do you expect the terminal to print?
- Fetch the current documentation metadata by running this command:
python source_probe.py
What does this command prove?
Python runs the probe against the live public repository. Each request reads the file metadata available on the repository's current main branch.
You should see three lines in the terminal. Each line shows a document title, a 12-character SHA prefix, and its repository path.
You now have live-source proof. Each run reads the current file metadata from GitHub.
The designed limitation is visible too. The probe has no chunking, search, persistence, or stored version history.
Probe request not completing?
- Confirm your Mac still has an internet connection.
- Confirm each configured path matches the corresponding entry in REMOTE_SOURCES.
- Confirm the request URL ends with ?ref=main.
- Help me troubleshoot my GitHub source probe.
✔️ Awesome, I've got everything!
Your probe now prints live version metadata for all three official documents.
ⓧ I'd like to double check the full code
from rag_core import REMOTE_SOURCES, fetch_github_source
for configured_source in REMOTE_SOURCES:
source = fetch_github_source(configured_source)
print(
f"{source['title']} | {source['source_version'][:12]} | "
f"{source['source_path']}"
)
Your project can now identify the current version of every tracked GitHub document. Next, you will turn those live files into a persistent searchable index.
Sync Changed Chunks into Qdrant
Your live source probe proved that GitHub exposes current documentation text. It also exposed a blob SHA for each tracked file.
Those results disappear when the probe stops. This step adds a persistent Qdrant vector database plus a version manifest that remembers every synchronized source.
In this step, get ready to:
- Clean each source before splitting it into heading-aware chunks.
- Embed each chunk before storing its vector with source metadata.
- Synchronize new, changed, unchanged, or retired sources through one repeatable job.
Clean and chunk every source
Raw Markdown includes frontmatter plus standalone template tags that add noise to retrieval. Heading-aware chunks preserve the section title that later appears in evidence and citations.
- In the open rag_core.py tab from earlier, select all existing content.
- Replace the selected content with the following sections in their displayed order.
- Start the file with the imports required for hashing, local paths, embeddings, and vector storage:
from base64 import b64decode
from hashlib import sha256
import json
from pathlib import Path
from urllib.parse import quote
from urllib.request import Request, urlopen
from openai import OpenAI
from qdrant_client import QdrantClient, models
What do these imports add?
- The sha256 helper creates stable versions for local content plus deterministic point IDs.
- The Path class reads the client overlay and writes the manifest.
- The OpenAI client creates embeddings. The QdrantClient client persists their vectors locally.
- Add the models, storage paths, source catalog, and reusable Qdrant client directly below the imports:
GENERATION_MODEL = "gpt-6-luna"
EMBEDDING_MODEL = "text-embedding-3-small"
COLLECTION_NAME = "github_actions_live_docs"
QDRANT_PATH = "qdrant_data"
MANIFEST_PATH = Path("manifest.json")
ASSUMPTIONS_PATH = Path("assumptions.md")
GITHUB_API_ROOT = "https://api.github.com/repos/github/docs/contents"
GITHUB_API_VERSION = "2026-03-10"
MAX_CHARS = 1800
REMOTE_SOURCES = [
{"path": "content/actions/reference/workflows-and-actions/workflow-syntax.md", "label": "Workflow syntax for GitHub Actions"},
{"path": "content/actions/reference/workflows-and-actions/workflow-commands.md", "label": "Workflow commands for GitHub Actions"},
{"path": "content/actions/reference/workflows-and-actions/dependency-caching.md", "label": "Dependency caching reference"},
]
_qdrant_client = None
def get_qdrant_client():
global _qdrant_client
if _qdrant_client is None:
_qdrant_client = QdrantClient(path=QDRANT_PATH)
return _qdrant_client
How does persistence start?
- The QDRANT_PATH value points the local client at qdrant_data so vectors survive between runs.
- The COLLECTION_NAME value gives this project a dedicated vector collection.
- The cached _qdrant_client reference keeps one local client active inside a process.
- Add the source normalization functions directly below get_qdrant_client():
def extract_title(text, fallback):
for line in text.splitlines()[:20]:
if line.startswith("title:"):
return line.split(":", 1)[1].strip().strip('"')
return fallback
def fetch_github_source(source):
encoded_path = quote(source["path"], safe="/")
request = Request(f"{GITHUB_API_ROOT}/{encoded_path}?ref=main", headers={"Accept": "application/vnd.github+json", "X-GitHub-Api-Version": GITHUB_API_VERSION, "User-Agent": "live-docs-rag-demo"})
with urlopen(request, timeout=30) as response:
payload = json.load(response)
text = b64decode(payload["content"]).decode("utf-8")
return {"source_path": payload["path"], "source_version": payload["sha"], "source_url": payload["html_url"], "source_kind": "official", "title": extract_title(text, source["label"]), "text": text}
def load_assumption_source():
text = ASSUMPTIONS_PATH.read_text(encoding="utf-8")
return {"source_path": str(ASSUMPTIONS_PATH), "source_version": sha256(text.encode("utf-8")).hexdigest(), "source_url": "", "source_kind": "assumption", "title": "Client Assumptions", "text": text}
How does the overlay stay distinct?
- Remote sources retain the blob SHA returned by GitHub as their source_version value.
- The local assumptions.md file receives a SHA-256 digest calculated from its text.
- The source_kind field labels the overlay as assumption so later answers can distinguish it from official content.
- Add the Markdown cleaning function directly below load_assumption_source():
def clean_markdown(text):
lines = text.splitlines()
if lines and lines[0].strip() == "---":
for index in range(1, len(lines)):
if lines[index].strip() == "---":
lines = lines[index + 1:]
break
return "\n".join(line for line in lines if not ((line.strip().startswith("{%") and line.strip().endswith("%}")) or (line.strip().startswith("{{") and line.strip().endswith("}}")))).strip()
What does the cleaner remove?
- The opening frontmatter block contains page metadata that does not belong in the retrieved documentation text.
- Standalone template tags depend on GitHub's documentation rendering system. Removing them reduces retrieval noise without trying to reproduce that rendering system.
- Add the heading-aware chunking function directly below clean_markdown():
def chunk_source(source):
chunks = []
for section_index, raw_section in enumerate(clean_markdown(source["text"]).split("\n## ")):
section = raw_section.strip()
if not section:
continue
if section_index > 0:
section = "## " + section
section_title = section.splitlines()[0].lstrip("# ").strip() or source["title"]
current = ""
for paragraph in section.split("\n\n"):
for piece in [paragraph.strip()[start:start + 1400] for start in range(0, len(paragraph.strip()), 1400)]:
if not piece:
continue
candidate = f"{current}\n\n{piece}".strip()
if current and len(candidate) > MAX_CHARS:
chunks.append({"title": section_title, "text": current})
current = piece
else:
current = candidate
if current:
chunks.append({"title": section_title, "text": current})
return chunks
How does chunking retain context?
- The function starts a section whenever it encounters a level-two Markdown heading.
- The section_title value follows every chunk created from that section.
- The MAX_CHARS boundary prevents a chunk from growing beyond 1,800 characters. Long paragraphs are divided into pieces of up to 1,400 characters first.
Embed chunks in persistent Qdrant
An embedding converts each chunk into numbers that represent its meaning. The OpenAI API creates those numbers before Qdrant stores them with searchable payload metadata.
- Add the embedding, manifest, point ID, and collection helpers directly below chunk_source():
def embed_texts(texts):
response = OpenAI().embeddings.create(model=EMBEDDING_MODEL, input=texts, encoding_format="float")
return [item.embedding for item in response.data]
def load_manifest():
return json.loads(MANIFEST_PATH.read_text(encoding="utf-8")) if MANIFEST_PATH.exists() else {}
def save_manifest(manifest):
MANIFEST_PATH.write_text(json.dumps(manifest, indent=2, sort_keys=True), encoding="utf-8")
def make_point_id(source_path, chunk_index):
return int(sha256(f"{source_path}:{chunk_index}".encode("utf-8")).hexdigest()[:15], 16)
def ensure_collection(client, vector_size):
if not client.collection_exists(COLLECTION_NAME):
client.create_collection(collection_name=COLLECTION_NAME, vectors_config=models.VectorParams(size=vector_size, distance=models.Distance.COSINE))
How do vectors become persistent?
- The embed_texts() function sends a batch of chunk text to text-embedding-3-small.
- The first returned embedding determines the collection's vector size.
- The make_point_id() function derives the same integer from a source path plus chunk position on every run.
- The manifest.json file records which point IDs belong to each source.
Synchronize versions and verify the index
The manifest turns synchronization into a comparison. Each run can skip a matching version or remove vectors whose source path has retired.
- Start the synchronization function directly below ensure_collection() with the source loading and retirement logic:
def sync_all():
client = get_qdrant_client()
manifest = load_manifest()
if manifest and not client.collection_exists(COLLECTION_NAME):
manifest = {}
sources = [fetch_github_source(source) for source in REMOTE_SOURCES]
sources.append(load_assumption_source())
active_paths = {source["source_path"] for source in sources}
stats = {"updated_sources": [], "skipped_sources": [], "removed_sources": [], "deleted_points": 0, "upserted_points": 0}
for retired_path in sorted(set(manifest) - active_paths):
retired = manifest.pop(retired_path)
if retired["point_ids"] and client.collection_exists(COLLECTION_NAME):
client.delete(collection_name=COLLECTION_NAME, points_selector=models.PointIdsList(points=retired["point_ids"]))
stats["deleted_points"] += len(retired["point_ids"])
stats["removed_sources"].append(retired_path)
How does retirement work?
- The active source paths come from the three configured remote files plus assumptions.md.
- A path found only in the manifest has retired from the active catalog.
- The function deletes that source's recorded point IDs before removing its manifest entry.
- Complete sync_all() by adding the version comparison and replacement logic directly after the retirement loop:
for source in sources:
source_path = source["source_path"]
previous = manifest.get(source_path)
if previous and previous["version"] == source["source_version"]:
stats["skipped_sources"].append(source_path)
continue
chunks = chunk_source(source)
embeddings = embed_texts([chunk["text"] for chunk in chunks])
ensure_collection(client, len(embeddings[0]))
if previous and previous["point_ids"]:
client.delete(collection_name=COLLECTION_NAME, points_selector=models.PointIdsList(points=previous["point_ids"]))
stats["deleted_points"] += len(previous["point_ids"])
point_ids = [make_point_id(source_path, index) for index in range(len(chunks))]
points = [models.PointStruct(id=point_id, vector=embedding, payload={"source_path": source_path, "source_version": source["source_version"], "source_url": source["source_url"], "source_kind": source["source_kind"], "title": chunk["title"], "text": chunk["text"]}) for point_id, chunk, embedding in zip(point_ids, chunks, embeddings)]
client.upsert(collection_name=COLLECTION_NAME, points=points)
manifest[source_path] = {"version": source["source_version"], "point_ids": point_ids, "source_kind": source["source_kind"], "title": source["title"]}
stats["updated_sources"].append(source_path)
stats["upserted_points"] += len(points)
save_manifest(manifest)
return stats
How does incremental replacement work?
- A matching version moves the source path into skipped_sources without creating new embeddings.
- A new or changed version produces fresh chunks plus embeddings.
- Existing point IDs are deleted before replacement points are upserted.
- Every point payload stores its path, version, URL, type, section title, and retrieved text.
- Save rag_core.py.
✔️ Awesome, I've got everything!
Your synchronization core now covers cleaning, chunking, embedding, persistence, version comparison, replacement, and retirement.
ⓧ I'd like to double check the full code
Compare your saved rag_core.py with this complete file.
from base64 import b64decode
from hashlib import sha256
import json
from pathlib import Path
from urllib.parse import quote
from urllib.request import Request, urlopen
from openai import OpenAI
from qdrant_client import QdrantClient, models
GENERATION_MODEL = "gpt-6-luna"
EMBEDDING_MODEL = "text-embedding-3-small"
COLLECTION_NAME = "github_actions_live_docs"
QDRANT_PATH = "qdrant_data"
MANIFEST_PATH = Path("manifest.json")
ASSUMPTIONS_PATH = Path("assumptions.md")
GITHUB_API_ROOT = "https://api.github.com/repos/github/docs/contents"
GITHUB_API_VERSION = "2026-03-10"
MAX_CHARS = 1800
REMOTE_SOURCES = [
{"path": "content/actions/reference/workflows-and-actions/workflow-syntax.md", "label": "Workflow syntax for GitHub Actions"},
{"path": "content/actions/reference/workflows-and-actions/workflow-commands.md", "label": "Workflow commands for GitHub Actions"},
{"path": "content/actions/reference/workflows-and-actions/dependency-caching.md", "label": "Dependency caching reference"},
]
_qdrant_client = None
def get_qdrant_client():
global _qdrant_client
if _qdrant_client is None:
_qdrant_client = QdrantClient(path=QDRANT_PATH)
return _qdrant_client
def extract_title(text, fallback):
for line in text.splitlines()[:20]:
if line.startswith("title:"):
return line.split(":", 1)[1].strip().strip('"')
return fallback
def fetch_github_source(source):
encoded_path = quote(source["path"], safe="/")
request = Request(f"{GITHUB_API_ROOT}/{encoded_path}?ref=main", headers={"Accept": "application/vnd.github+json", "X-GitHub-Api-Version": GITHUB_API_VERSION, "User-Agent": "live-docs-rag-demo"})
with urlopen(request, timeout=30) as response:
payload = json.load(response)
text = b64decode(payload["content"]).decode("utf-8")
return {"source_path": payload["path"], "source_version": payload["sha"], "source_url": payload["html_url"], "source_kind": "official", "title": extract_title(text, source["label"]), "text": text}
def load_assumption_source():
text = ASSUMPTIONS_PATH.read_text(encoding="utf-8")
return {"source_path": str(ASSUMPTIONS_PATH), "source_version": sha256(text.encode("utf-8")).hexdigest(), "source_url": "", "source_kind": "assumption", "title": "Client Assumptions", "text": text}
def clean_markdown(text):
lines = text.splitlines()
if lines and lines[0].strip() == "---":
for index in range(1, len(lines)):
if lines[index].strip() == "---":
lines = lines[index + 1:]
break
return "\n".join(line for line in lines if not ((line.strip().startswith("{%") and line.strip().endswith("%}")) or (line.strip().startswith("{{") and line.strip().endswith("}}")))).strip()
def chunk_source(source):
chunks = []
for section_index, raw_section in enumerate(clean_markdown(source["text"]).split("\n## ")):
section = raw_section.strip()
if not section:
continue
if section_index > 0:
section = "## " + section
section_title = section.splitlines()[0].lstrip("# ").strip() or source["title"]
current = ""
for paragraph in section.split("\n\n"):
for piece in [paragraph.strip()[start:start + 1400] for start in range(0, len(paragraph.strip()), 1400)]:
if not piece:
continue
candidate = f"{current}\n\n{piece}".strip()
if current and len(candidate) > MAX_CHARS:
chunks.append({"title": section_title, "text": current})
current = piece
else:
current = candidate
if current:
chunks.append({"title": section_title, "text": current})
return chunks
def embed_texts(texts):
response = OpenAI().embeddings.create(model=EMBEDDING_MODEL, input=texts, encoding_format="float")
return [item.embedding for item in response.data]
def load_manifest():
return json.loads(MANIFEST_PATH.read_text(encoding="utf-8")) if MANIFEST_PATH.exists() else {}
def save_manifest(manifest):
MANIFEST_PATH.write_text(json.dumps(manifest, indent=2, sort_keys=True), encoding="utf-8")
def make_point_id(source_path, chunk_index):
return int(sha256(f"{source_path}:{chunk_index}".encode("utf-8")).hexdigest()[:15], 16)
def ensure_collection(client, vector_size):
if not client.collection_exists(COLLECTION_NAME):
client.create_collection(collection_name=COLLECTION_NAME, vectors_config=models.VectorParams(size=vector_size, distance=models.Distance.COSINE))
def sync_all():
client = get_qdrant_client()
manifest = load_manifest()
if manifest and not client.collection_exists(COLLECTION_NAME):
manifest = {}
sources = [fetch_github_source(source) for source in REMOTE_SOURCES]
sources.append(load_assumption_source())
active_paths = {source["source_path"] for source in sources}
stats = {"updated_sources": [], "skipped_sources": [], "removed_sources": [], "deleted_points": 0, "upserted_points": 0}
for retired_path in sorted(set(manifest) - active_paths):
retired = manifest.pop(retired_path)
if retired["point_ids"] and client.collection_exists(COLLECTION_NAME):
client.delete(collection_name=COLLECTION_NAME, points_selector=models.PointIdsList(points=retired["point_ids"]))
stats["deleted_points"] += len(retired["point_ids"])
stats["removed_sources"].append(retired_path)
for source in sources:
source_path = source["source_path"]
previous = manifest.get(source_path)
if previous and previous["version"] == source["source_version"]:
stats["skipped_sources"].append(source_path)
continue
chunks = chunk_source(source)
embeddings = embed_texts([chunk["text"] for chunk in chunks])
ensure_collection(client, len(embeddings[0]))
if previous and previous["point_ids"]:
client.delete(collection_name=COLLECTION_NAME, points_selector=models.PointIdsList(points=previous["point_ids"]))
stats["deleted_points"] += len(previous["point_ids"])
point_ids = [make_point_id(source_path, index) for index in range(len(chunks))]
points = [models.PointStruct(id=point_id, vector=embedding, payload={"source_path": source_path, "source_version": source["source_version"], "source_url": source["source_url"], "source_kind": source["source_kind"], "title": chunk["title"], "text": chunk["text"]}) for point_id, chunk, embedding in zip(point_ids, chunks, embeddings)]
client.upsert(collection_name=COLLECTION_NAME, points=points)
manifest[source_path] = {"version": source["source_version"], "point_ids": point_ids, "source_kind": source["source_kind"], "title": source["title"]}
stats["updated_sources"].append(source_path)
stats["upserted_points"] += len(points)
save_manifest(manifest)
return stats
The synchronization logic needs a small command-line entry point. This keeps manual sync runs separate from the reusable functions in rag_core.py.
- In the VS Code Explorer sidebar, create sync_docs.py inside the same folder as rag_core.py.
- Paste this code into sync_docs.py:
import json
from rag_core import sync_all
print(json.dumps(sync_all(), indent=2))
What does this entry point do?
The script calls sync_all() once. It prints the returned statistics as formatted JSON so each lifecycle action is visible.
- Save sync_docs.py.
✔️ Awesome, I've got everything!
Your command-line sync entry point is ready.
ⓧ I'd like to double check the full code
Compare your saved sync_docs.py with this complete file.
import json
from rag_core import sync_all
print(json.dumps(sync_all(), indent=2))
Your API budget stays protected
This first sync sends chunk text to the OpenAI API for embedding. The hard spend limit you configured earlier remains your budget boundary.
Before you run the sync, which report fields do you expect to change when every source is new?
- Stop the Streamlit Hello process with Ctrl+C if it is still using your terminal.
- Run the first synchronization from the activated environment with this command:
python sync_docs.py
The first run fetches three live GitHub files plus the local overlay. It also creates embeddings for every resulting chunk.
You will see four paths under updated_sources and a value above zero for upserted_points.
- Return to the VS Code Explorer sidebar from earlier.
- Confirm that manifest.json now appears beside rag_core.py.
- Confirm that the qdrant_data directory now appears in the same folder.
Sync did not complete?
- Confirm that the activated environment still appears in your terminal prompt. The pinned packages are installed inside that environment.
- Confirm that your OpenAI API key remains exported in the same terminal session.
- Check that assumptions.md sits beside rag_core.py because the sync reads it through a relative path.
- Ask for help with the exact terminal output by using Help me troubleshoot my first Qdrant synchronization run..
That is the storage layer working. Your four sources now survive between runs with their versions, vectors, metadata, and point ownership recorded.
Next, you will query this synchronized snapshot and turn its evidence into a source-aware answer.
Ask the Live Documentation Assistant
Your persistent Qdrant collection now holds synchronized chunks from the official GitHub documentation plus the client overlay. Those vectors become useful when a question can retrieve evidence from the same stored snapshot.
In this step, you will add semantic retrieval to rag_core.py. You will connect it to the OpenAI Responses API. You will present the result through a Streamlit chat interface.
In this step, get ready to:
- Retrieve the four closest chunks for each question.
- Generate answers that distinguish official documentation from client assumptions.
- Build a browser interface with sync status plus expandable evidence.
Retrieve evidence from Qdrant
A retrieval query converts the question into an embedding. Qdrant compares that vector with the stored documentation vectors to find the closest evidence.
- Switch back to rag_core.py in Visual Studio Code.
- Scroll below the existing sync_all() function.
- Add the collection check plus manifest table helpers by pasting this code:
def collection_ready():
return get_qdrant_client().collection_exists(COLLECTION_NAME)
def manifest_rows():
rows = []
for path, entry in sorted(load_manifest().items()):
rows.append(
{
"Source": entry["title"],
"Type": entry["source_kind"],
"Version": entry["version"][:12],
"Chunks": len(entry["point_ids"]),
"Path": path,
}
)
return rows
What Do These Helpers Provide?
- The collection_ready() helper confirms that the persistent collection exists before the chat performs retrieval.
- The manifest_rows() helper converts the manifest dictionary into dashboard rows.
- Each dashboard row exposes the source title plus its type.
- The version prefix plus chunk count makes the synchronized state visible.
- Save rag_core.py.
- Confirm the existing sync still imports the updated file by running:
python sync_docs.py
What Does This Check Prove?
The command imports the updated rag_core.py before running the synchronization logic. You should see unchanged paths under skipped_sources with no syntax error.
Does the Sync Stop Before the Report?
- Check that both new functions start at the far-left edge of rag_core.py.
- Check the closing brackets inside the dictionary returned by manifest_rows().
- Use this guided prompt if the import still fails: Help me find the syntax issue in the collection_ready() and manifest_rows() functions I added to rag_core.py.
The next helper performs the actual similarity search. It keeps the payload attached so the interface can show where every retrieved chunk came from.
- Add the retrieval helper directly below manifest_rows() by pasting this code:
def query_docs(question, limit=4):
if not collection_ready():
return []
query_vector = embed_texts([question])[0]
response = get_qdrant_client().query_points(
collection_name=COLLECTION_NAME,
query=query_vector,
limit=limit,
with_payload=True,
)
results = []
for point in response.points:
payload = point.payload or {}
results.append(
{
"score": point.score,
"title": payload.get("title", "Untitled"),
"text": payload.get("text", ""),
"source_path": payload.get("source_path", ""),
"source_version": payload.get("source_version", ""),
"source_url": payload.get("source_url", ""),
"source_kind": payload.get("source_kind", "unknown"),
}
)
return results
How Does Retrieval Work?
- The question passes through embed_texts() so it uses the same embedding model as the stored chunks.
- The query_points() call requests the four closest points by default.
- The with_payload=True setting returns the metadata required for labels plus citations.
- The result dictionaries preserve the similarity score plus the full source record.
- Save rag_core.py.
- Confirm the retrieval helper imports successfully by running:
python sync_docs.py
What Should the Report Show?
You should see the synchronization report complete without rebuilding unchanged vectors. That result confirms the retrieval code has not broken the existing sync path.
Seeing a Python Syntax Error?
- Check that def query_docs(question, limit=4): starts below the completed manifest_rows() function.
- Check that every key inside the result dictionary has a matching comma.
- Use this guided prompt for a focused comparison: Help me compare my query_docs() function with the expected indentation and brackets.
Generate source-aware answers
Retrieved chunks need explicit labels before they reach the generation model. These labels keep official GitHub documentation separate from private client guidance inside the prompt.
- Add the first half of answer_question() directly below query_docs() by pasting this code:
def answer_question(question):
sources = query_docs(question)
if not sources:
return "Run the documentation sync before asking questions.", []
context_blocks = []
for source in sources:
label = (
"Official GitHub Docs"
if source["source_kind"] == "official"
else "Client assumption"
)
context_blocks.append(
f"[{label} | Source: {source['title']}]\n{source['text']}"
)
Why Label the Context?
- The function first retrieves evidence through query_docs().
- An empty result returns a clear instruction to synchronize the collection.
- Each retrieved chunk receives either the Official GitHub Docs label or the Client assumption label.
- The source title travels with the evidence into the generation prompt.
- Continue the same answer_question() function by pasting this indented code directly below the first half:
response = OpenAI().responses.create(
model=GENERATION_MODEL,
reasoning={"effort": "none"},
instructions=(
"Answer only from the supplied synchronized sources. "
"Treat Official GitHub Docs as vendor documentation. "
"Treat Client assumption content as private guidance and label it clearly. "
"If the sources do not support an answer, say that the synchronized sources "
"do not contain enough information. End factual sentences with citations in "
"the format [Source: section title]."
),
input=(
"Synchronized sources:\n\n"
+ "\n\n".join(context_blocks)
+ f"\n\nQuestion: {question}"
),
)
return response.output_text, sources
How Is the Answer Constrained?
- The request uses the existing GENERATION_MODEL value.
- The instructions restrict answers to the synchronized evidence.
- Unsupported questions receive an explicit insufficiency response.
- The function returns the answer plus the retrieved records so the interface can display its evidence.
- Save rag_core.py.
- Confirm the completed answer function imports with the existing pipeline by running:
python sync_docs.py
What Does a Successful Import Confirm?
The JSON report should print after all four new helpers load. Unchanged source versions should remain skipped.
Does the Responses Call Fail to Import?
- Check that the second snippet remains indented inside answer_question().
- Check that return response.output_text, sources aligns with the response = OpenAI().responses.create( line.
- Use this guided prompt if the function still fails: Help me repair the indentation in answer_question() without changing its source-labeling instructions.
✔️ Awesome, I've got everything!
Great. Save rag_core.py before you build the browser interface.
ⓧ I'd like to double check the full code
Compare your complete rag_core.py with this reference.
from base64 import b64decode
from hashlib import sha256
import json
from pathlib import Path
from urllib.parse import quote
from urllib.request import Request, urlopen
from openai import OpenAI
from qdrant_client import QdrantClient, models
GENERATION_MODEL = "gpt-6-luna"
EMBEDDING_MODEL = "text-embedding-3-small"
COLLECTION_NAME = "github_actions_live_docs"
QDRANT_PATH = "qdrant_data"
MANIFEST_PATH = Path("manifest.json")
ASSUMPTIONS_PATH = Path("assumptions.md")
GITHUB_API_ROOT = "https://api.github.com/repos/github/docs/contents"
GITHUB_API_VERSION = "2026-03-10"
MAX_CHARS = 1800
REMOTE_SOURCES = [
{
"path": "content/actions/reference/workflows-and-actions/workflow-syntax.md",
"label": "Workflow syntax for GitHub Actions",
},
{
"path": "content/actions/reference/workflows-and-actions/workflow-commands.md",
"label": "Workflow commands for GitHub Actions",
},
{
"path": "content/actions/reference/workflows-and-actions/dependency-caching.md",
"label": "Dependency caching reference",
},
]
_qdrant_client = None
def get_qdrant_client():
global _qdrant_client
if _qdrant_client is None:
_qdrant_client = QdrantClient(path=QDRANT_PATH)
return _qdrant_client
def extract_title(text, fallback):
for line in text.splitlines()[:20]:
if line.startswith("title:"):
return line.split(":", 1)[1].strip().strip('"')
return fallback
def fetch_github_source(source):
encoded_path = quote(source["path"], safe="/")
url = f"{GITHUB_API_ROOT}/{encoded_path}?ref=main"
request = Request(
url,
headers={
"Accept": "application/vnd.github+json",
"X-GitHub-Api-Version": GITHUB_API_VERSION,
"User-Agent": "live-docs-rag-demo",
},
)
with urlopen(request, timeout=30) as response:
payload = json.load(response)
text = b64decode(payload["content"]).decode("utf-8")
return {
"source_path": payload["path"],
"source_version": payload["sha"],
"source_url": payload["html_url"],
"source_kind": "official",
"title": extract_title(text, source["label"]),
"text": text,
}
def load_assumption_source():
text = ASSUMPTIONS_PATH.read_text(encoding="utf-8")
return {
"source_path": str(ASSUMPTIONS_PATH),
"source_version": sha256(text.encode("utf-8")).hexdigest(),
"source_url": "",
"source_kind": "assumption",
"title": "Client Assumptions",
"text": text,
}
def clean_markdown(text):
lines = text.splitlines()
if lines and lines[0].strip() == "---":
for index in range(1, len(lines)):
if lines[index].strip() == "---":
lines = lines[index + 1 :]
break
cleaned = []
for line in lines:
stripped = line.strip()
if stripped.startswith("{%") and stripped.endswith("%}"):
continue
if stripped.startswith("{{") and stripped.endswith("}}"):
continue
cleaned.append(line)
return "\n".join(cleaned).strip()
def chunk_source(source):
text = clean_markdown(source["text"])
raw_sections = text.split("\n## ")
chunks = []
for section_index, raw_section in enumerate(raw_sections):
section = raw_section.strip()
if not section:
continue
if section_index > 0:
section = "## " + section
first_line = section.splitlines()[0]
section_title = first_line.lstrip("# ").strip() or source["title"]
current = ""
for paragraph in section.split("\n\n"):
paragraph = paragraph.strip()
if not paragraph:
continue
pieces = [
paragraph[start : start + 1400]
for start in range(0, len(paragraph), 1400)
]
for piece in pieces:
candidate = f"{current}\n\n{piece}".strip()
if current and len(candidate) > MAX_CHARS:
chunks.append({"title": section_title, "text": current})
current = piece
else:
current = candidate
if current:
chunks.append({"title": section_title, "text": current})
return chunks
def embed_texts(texts):
response = OpenAI().embeddings.create(
model=EMBEDDING_MODEL,
input=texts,
encoding_format="float",
)
return [item.embedding for item in response.data]
def load_manifest():
if not MANIFEST_PATH.exists():
return {}
return json.loads(MANIFEST_PATH.read_text(encoding="utf-8"))
def save_manifest(manifest):
MANIFEST_PATH.write_text(
json.dumps(manifest, indent=2, sort_keys=True),
encoding="utf-8",
)
def make_point_id(source_path, chunk_index):
raw = f"{source_path}:{chunk_index}".encode("utf-8")
return int(sha256(raw).hexdigest()[:15], 16)
def ensure_collection(client, vector_size):
if not client.collection_exists(COLLECTION_NAME):
client.create_collection(
collection_name=COLLECTION_NAME,
vectors_config=models.VectorParams(
size=vector_size,
distance=models.Distance.COSINE,
),
)
def sync_all():
client = get_qdrant_client()
manifest = load_manifest()
if manifest and not client.collection_exists(COLLECTION_NAME):
manifest = {}
sources = [fetch_github_source(source) for source in REMOTE_SOURCES]
sources.append(load_assumption_source())
active_paths = {source["source_path"] for source in sources}
stats = {
"updated_sources": [],
"skipped_sources": [],
"removed_sources": [],
"deleted_points": 0,
"upserted_points": 0,
}
for retired_path in sorted(set(manifest) - active_paths):
retired = manifest.pop(retired_path)
if retired["point_ids"] and client.collection_exists(COLLECTION_NAME):
client.delete(
collection_name=COLLECTION_NAME,
points_selector=models.PointIdsList(points=retired["point_ids"]),
)
stats["deleted_points"] += len(retired["point_ids"])
stats["removed_sources"].append(retired_path)
for source in sources:
source_path = source["source_path"]
previous = manifest.get(source_path)
if previous and previous["version"] == source["source_version"]:
stats["skipped_sources"].append(source_path)
continue
chunks = chunk_source(source)
embeddings = embed_texts([chunk["text"] for chunk in chunks])
ensure_collection(client, len(embeddings[0]))
if previous and previous["point_ids"]:
client.delete(
collection_name=COLLECTION_NAME,
points_selector=models.PointIdsList(points=previous["point_ids"]),
)
stats["deleted_points"] += len(previous["point_ids"])
point_ids = [
make_point_id(source_path, index) for index in range(len(chunks))
]
points = []
for point_id, chunk, embedding in zip(point_ids, chunks, embeddings):
points.append(
models.PointStruct(
id=point_id,
vector=embedding,
payload={
"source_path": source_path,
"source_version": source["source_version"],
"source_url": source["source_url"],
"source_kind": source["source_kind"],
"title": chunk["title"],
"text": chunk["text"],
},
)
)
client.upsert(collection_name=COLLECTION_NAME, points=points)
manifest[source_path] = {
"version": source["source_version"],
"point_ids": point_ids,
"source_kind": source["source_kind"],
"title": source["title"],
}
stats["updated_sources"].append(source_path)
stats["upserted_points"] += len(points)
save_manifest(manifest)
return stats
def collection_ready():
return get_qdrant_client().collection_exists(COLLECTION_NAME)
def manifest_rows():
rows = []
for path, entry in sorted(load_manifest().items()):
rows.append(
{
"Source": entry["title"],
"Type": entry["source_kind"],
"Version": entry["version"][:12],
"Chunks": len(entry["point_ids"]),
"Path": path,
}
)
return rows
def query_docs(question, limit=4):
if not collection_ready():
return []
query_vector = embed_texts([question])[0]
response = get_qdrant_client().query_points(
collection_name=COLLECTION_NAME,
query=query_vector,
limit=limit,
with_payload=True,
)
results = []
for point in response.points:
payload = point.payload or {}
results.append(
{
"score": point.score,
"title": payload.get("title", "Untitled"),
"text": payload.get("text", ""),
"source_path": payload.get("source_path", ""),
"source_version": payload.get("source_version", ""),
"source_url": payload.get("source_url", ""),
"source_kind": payload.get("source_kind", "unknown"),
}
)
return results
def answer_question(question):
sources = query_docs(question)
if not sources:
return "Run the documentation sync before asking questions.", []
context_blocks = []
for source in sources:
label = (
"Official GitHub Docs"
if source["source_kind"] == "official"
else "Client assumption"
)
context_blocks.append(
f"[{label} | Source: {source['title']}]\n{source['text']}"
)
response = OpenAI().responses.create(
model=GENERATION_MODEL,
reasoning={"effort": "none"},
instructions=(
"Answer only from the supplied synchronized sources. "
"Treat Official GitHub Docs as vendor documentation. "
"Treat Client assumption content as private guidance and label it clearly. "
"If the sources do not support an answer, say that the synchronized sources "
"do not contain enough information. End factual sentences with citations in "
"the format [Source: section title]."
),
input=(
"Synchronized sources:\n\n"
+ "\n\n".join(context_blocks)
+ f"\n\nQuestion: {question}"
),
)
return response.output_text, sources
What Should You Compare?
Confirm that the four new helpers appear after sync_all(). Their identifiers plus model settings must match this reference exactly.
Build and test the Streamlit chat
The interface combines synchronization controls with the current manifest state. It also keeps each chat response beside the evidence that produced it.
- Create app.py inside the same project folder as rag_core.py using the Visual Studio Code Explorer.
- Add the dashboard imports plus synchronized-state panel by pasting this code:
import streamlit as st
from rag_core import (
answer_question,
collection_ready,
manifest_rows,
sync_all,
)
st.write("# Self-Updating GitHub Actions Docs RAG")
st.write(
"The index combines live official GitHub Docs with an explicitly labeled "
"client-assumption overlay."
)
if st.button("Check live docs and sync Qdrant", type="primary"):
st.session_state["sync_stats"] = sync_all()
if "sync_stats" in st.session_state:
st.write("## Latest sync report")
st.write(st.session_state["sync_stats"])
rows = manifest_rows()
st.metric("Tracked sources", len(rows))
st.metric("Stored chunks", sum(row["Chunks"] for row in rows))
if rows:
st.dataframe(rows, hide_index=True)
What Does the Dashboard Show?
- The primary button runs the same sync_all() lifecycle used by the terminal entry point.
- Session state keeps the latest sync report available across interface reruns.
- The metrics summarize the current manifest plus its stored point ownership.
- The table exposes each tracked source with its type plus version prefix.
The Streamlit server keeps this terminal process occupied while the app is running. Leave it active while you complete the remaining interface edits.
- Save app.py.
Before you launch the dashboard, which manifest details do you expect the page to expose?
- Launch the dashboard from the activated environment by running:
python -m streamlit run app.py
What Does This Command Start?
Streamlit runs app.py as a local web application. You should see the dashboard in a browser tab.
You should see four tracked sources plus a nonzero stored-chunk count. The table should list three official sources plus Client Assumptions.
Dashboard Not Opening?
- Confirm that the terminal prompt shows the activated virtual environment.
- Check that app.py sits beside rag_core.py.
- Check the terminal for the first referenced line in app.py or rag_core.py.
- Use this guided prompt for environment-specific help: Help me diagnose why my Streamlit app cannot import rag_core or open the local dashboard.
Chat history belongs in session state because each submitted message reruns the app script. Rendering stored messages rebuilds the conversation after every question.
- Switch back to app.py from earlier.
- Add the chat heading plus history renderer below the manifest table by pasting this code:
st.write("---")
st.write("## Ask the synchronized snapshot")
if "messages" not in st.session_state:
st.session_state["messages"] = []
for message in st.session_state["messages"]:
with st.chat_message(message["role"]):
st.write(message["content"])
if message.get("sources"):
with st.expander("Retrieved evidence"):
for source in message["sources"]:
st.write(
f"**{source['title']}** | {source['source_kind']} | "
f"score {source['score']:.3f} | "
f"version {source['source_version'][:12]}"
)
if source["source_url"]:
st.write(f"[Open official source]({source['source_url']})")
st.write(source["text"])
How Is Conversation History Rebuilt?
- The messages list stores the conversation for the browser session.
- Each stored role selects the matching chat-message container.
- Assistant records can include the retrieved source list.
- The evidence expander displays the source type plus score plus version prefix plus chunk text.
- Save app.py.
- Return to the running browser tab.
You should now see the Ask the synchronized snapshot heading below the manifest table.
Chat Heading Missing?
- Confirm that the new block starts below the completed if rows: block.
- Check that the lines inside both with blocks remain indented.
- Use this guided prompt if the page reports an indentation problem: Help me fix the chat-history block in app.py without changing its retrieved-evidence fields.
The input handler sends a question through the collection check plus answer pipeline. It also stores the user's message before the assistant response is rendered.
- Add the question input plus answer selection below the history loop by pasting this code:
if prompt := st.chat_input(
"Ask about GitHub Actions workflows, commands, caching, or client assumptions"
):
st.session_state["messages"].append({"role": "user", "content": prompt})
with st.chat_message("user"):
st.write(prompt)
if collection_ready():
answer, sources = answer_question(prompt)
else:
answer, sources = "Select the sync button before asking questions.", []
What Happens When You Submit a Question?
- The chat input returns the submitted question as prompt.
- The user's message enters session history immediately.
- A ready collection sends the question through answer_question().
- A missing collection produces a sync instruction without attempting retrieval.
- Save app.py.
- Return to the running browser tab.
You should see a chat input field beneath the synchronized-snapshot heading.
- Complete the same input block by pasting this indented response renderer directly below the else branch:
with st.chat_message("assistant"):
st.write(answer)
if sources:
with st.expander("Retrieved evidence"):
for source in sources:
st.write(
f"**{source['title']}** | {source['source_kind']} | "
f"score {source['score']:.3f} | "
f"version {source['source_version'][:12]}"
)
if source["source_url"]:
st.write(f"[Open official source]({source['source_url']})")
st.write(source["text"])
st.session_state["messages"].append(
{"role": "assistant", "content": answer, "sources": sources}
)
How Is the Answer Presented?
- The assistant container displays the generated answer.
- Retrieved records appear inside an expandable evidence panel.
- Official records include a browser link from their stored payload.
- The completed assistant record enters session history with its evidence.
- Save app.py.
✔️ Awesome, I've got everything!
Your dashboard plus chat path is complete. Keep the Streamlit process running for the final check.
ⓧ I'd like to double check the full code
Compare your complete app.py with this reference.
import streamlit as st
from rag_core import (
answer_question,
collection_ready,
manifest_rows,
sync_all,
)
st.write("# Self-Updating GitHub Actions Docs RAG")
st.write(
"The index combines live official GitHub Docs with an explicitly labeled "
"client-assumption overlay."
)
if st.button("Check live docs and sync Qdrant", type="primary"):
st.session_state["sync_stats"] = sync_all()
if "sync_stats" in st.session_state:
st.write("## Latest sync report")
st.write(st.session_state["sync_stats"])
rows = manifest_rows()
st.metric("Tracked sources", len(rows))
st.metric("Stored chunks", sum(row["Chunks"] for row in rows))
if rows:
st.dataframe(rows, hide_index=True)
st.write("---")
st.write("## Ask the synchronized snapshot")
if "messages" not in st.session_state:
st.session_state["messages"] = []
for message in st.session_state["messages"]:
with st.chat_message(message["role"]):
st.write(message["content"])
if message.get("sources"):
with st.expander("Retrieved evidence"):
for source in message["sources"]:
st.write(
f"**{source['title']}** | {source['source_kind']} | "
f"score {source['score']:.3f} | "
f"version {source['source_version'][:12]}"
)
if source["source_url"]:
st.write(f"[Open official source]({source['source_url']})")
st.write(source["text"])
if prompt := st.chat_input(
"Ask about GitHub Actions workflows, commands, caching, or client assumptions"
):
st.session_state["messages"].append({"role": "user", "content": prompt})
with st.chat_message("user"):
st.write(prompt)
if collection_ready():
answer, sources = answer_question(prompt)
else:
answer, sources = "Select the sync button before asking questions.", []
with st.chat_message("assistant"):
st.write(answer)
if sources:
with st.expander("Retrieved evidence"):
for source in sources:
st.write(
f"**{source['title']}** | {source['source_kind']} | "
f"score {source['score']:.3f} | "
f"version {source['source_version'][:12]}"
)
if source["source_url"]:
st.write(f"[Open official source]({source['source_url']})")
st.write(source["text"])
st.session_state["messages"].append(
{"role": "assistant", "content": answer, "sources": sources}
)
What Should Match?
Check the order of the dashboard plus history plus input sections. Confirm that both evidence renderers show the same payload fields.
Before you ask, which source type do you expect to support a question about workflow-file placement?
- Return to the running Streamlit browser tab.
- Enter Where must a workflow file be stored? in the chat input.
- Submit the question.
- Expand Retrieved evidence below the answer.
You should see a concise cited answer supported by Workflow syntax for GitHub Actions. The evidence should identify the record as official plus show its score plus version prefix plus official source link plus chunk text.
That is the retrieval loop working from end to end. Your assistant can now turn the synchronized vector snapshot into a visible answer with inspectable evidence.
No Answer or Evidence?
- Confirm that the tracked-source metric shows four sources.
- Confirm that the stored-chunk metric is greater than zero.
- Confirm that the active shell still contains OPENAI_API_KEY.
- Use this guided prompt for the full retrieval path: Help me trace a Streamlit question through query_docs() and answer_question() to find why no evidence appears.
Your synchronized documentation is now searchable through a source-aware chat. Next, you will edit the client overlay to prove that one changed source can be replaced without rebuilding the unchanged GitHub content.
Prove Incremental Updates with a Client Overlay
The previous step gave you a source-aware Streamlit assistant backed by Qdrant. Now you will prove that its index updates incrementally.
Live GitHub files may stay unchanged during a short demo. A controlled edit to assumptions.md creates a change you can verify on demand.
In this step, get ready to:
- Capture the assistant's current answer about the client freshness target.
- Change only the client overlay and synchronize its replacement vectors.
- Restore the 15-minute target and verify the final source-aware answer.
Capture the current overlay answer
A baseline gives you a clear before-and-after comparison. The current answer should reflect the overlay version already stored in the collection.
Before you ask, which source type do you expect the assistant to use for a client-specific refresh target?
- Return to the Streamlit browser tab from the previous step.
- Enter How often should private guidance and vendor documentation be checked? in the chat field.
- Press Enter.
- Expand Retrieved evidence beneath the response.
You should see an answer based on a 15-minute target. The answer should identify the guidance as a Client assumption.
Good. You now have a baseline that makes the upcoming replacement visible.
Change only the client overlay
The local overlay uses a SHA-256 content hash as its source version. Changing one value produces a new version for the synchronization logic to detect.
- Return to Visual Studio Code from earlier.
- Select assumptions.md in the Explorer sidebar.
- Locate the ## Freshness target section.
- Replace 15 with 30 in the freshness statement.
- Save assumptions.md.
This sync makes a small paid embedding call through the OpenAI API. Your $10 spend control from Step 1 remains the budget guardrail.
Before you sync, which source do you expect to be updated?
- Return to the Streamlit browser tab.
- Click Check live docs and sync Qdrant.
How to read the sync report
- The updated_sources list contains assumptions.md.
- The skipped_sources list contains content/actions/reference/workflows-and-actions/workflow-syntax.md.
- The skipped_sources list contains content/actions/reference/workflows-and-actions/workflow-commands.md.
- The skipped_sources list contains content/actions/reference/workflows-and-actions/dependency-caching.md.
- The deleted_points value is greater than zero because the previous overlay points were removed.
- The upserted_points value is greater than zero because replacement overlay chunks were stored.
Overlay listed as skipped?
Confirm that assumptions.md is saved with the temporary 30-minute value. An unsaved edit leaves the previous content hash unchanged.
Select the sync button again after saving the file.
Help me understand why my changed assumptions.md file appears in skipped_sources.
The replacement vectors are now searchable. Asking the same question tests whether retrieval uses the new overlay snapshot.
Before you ask again, do you expect the answer to keep the old target or use the temporary target?
- Enter How often should private guidance and vendor documentation be checked? in the chat field again.
- Press Enter.
- Expand Retrieved evidence beneath the new response.
You should see the temporary 30-minute target in the new answer. The response should still label the guidance as a Client assumption.
You have proved that one changed overlay can be replaced without rebuilding the three official sources.
Restore the published target and verify the replacement
A controlled update demonstration should finish with the project content restored. The published overlay uses the 15-minute target from the completed project.
- Switch back to assumptions.md in Visual Studio Code.
- Replace 30 with 15 in the freshness statement.
✔️ Awesome, I've got everything!
Your overlay now matches the published file with its 15-minute freshness target.
ⓧ I'd like to double check the full code
- Compare your complete assumptions.md file with this reference:
# Client Assumptions
## Freshness target
This portfolio demo assumes the client wants the private guidance and vendor documentation checked every 15 minutes by an external scheduler.
## Source policy
Official GitHub documentation is authoritative for GitHub Actions behavior. This client overlay may add internal recommendations, but the assistant must label them as assumptions and must not present them as GitHub policy.
## Unsupported questions
When retrieved evidence does not support an answer, the assistant should say that the synchronized sources do not contain enough information.
- Save assumptions.md.
Before the final sync, which manifest entry do you expect to receive another new version?
- Return to the Streamlit browser tab.
- Click Check live docs and sync Qdrant.
Final sync checkpoint
- The three official GitHub paths remain in skipped_sources.
- The assumptions.md path appears in updated_sources.
- The deleted_points value counts the temporary overlay points that were removed.
- The upserted_points value counts the restored overlay chunks.
- The manifest table shows a new version prefix for Client Assumptions.
Before the final question, which freshness target should the synchronized snapshot return now?
- Enter How often should private guidance and vendor documentation be checked? in the chat field once more.
- Press Enter.
- Expand Retrieved evidence beneath the final response.
You should see the restored 15-minute target. The answer should identify it as a Client assumption.
You have now proved the whole incremental update loop with a controlled change.
Still seeing the temporary target?
Confirm that the latest sync report lists assumptions.md under updated_sources. A skipped overlay means the saved file still matches the manifest version.
Submit the question as a new chat message after the final sync.
Help me troubleshoot why my assistant still retrieves the temporary client freshness target.
Secret mission
Run a Source Retirement Drill
Temporarily retire one live GitHub Actions source. Watch the synchronizer remove its stale vectors. Restore the source to prove the index converges to the active catalog in both directions.
Clean Up Your Resources
Clean Up Your Resources
Your local Streamlit process and Qdrant database have no ongoing service cost. Choose whether to keep the project ready, pause it before returning, or delete its local artifacts and OpenAI API key.
Cost warning
OpenAI API usage is token-metered when you sync sources or ask the assistant a question. Keep the $10 hard spend limit enabled if you retain the key.
Hard-limit enforcement is not instantaneous. Recorded spend can slightly exceed the configured amount.
Resources you used:
- A running Streamlit process in your active terminal.
- An active OPENAI_API_KEY shell session.
- A local project folder containing PLAN.md, requirements.txt, .gitignore, assumptions.md, rag_core.py, source_probe.py, sync_docs.py, and app.py.
- A local .venv/ virtual environment.
- A persistent qdrant_data/ vector database.
- A generated manifest.json source-version record.
- An OpenAI API key created for this project.
Keep everything running
No action needed. Choose this if you plan to continue testing synchronization or want to demonstrate the assistant again soon.
- Keep the Streamlit app running while you actively use the assistant.
- Retain the .venv/ environment with its pinned packages.
- Retain qdrant_data/ with vectors for the three official sources and assumptions.md.
- Retain manifest.json with its four restored source entries.
- Keep the $10 hard spend limit enabled for future API activity.
- Use the Check live docs and sync Qdrant button for future syncs while Streamlit is running.
- Avoid starting the standalone sync_docs.py sync while Streamlit has the local Qdrant database open.
Pause - I'll come back to this later
Shut down the running process to free local memory. Your project files and synchronized snapshot stay available for later.
- Stop Streamlit by pressing Ctrl+C in the terminal.
- Close the terminal window to end the active OPENAI_API_KEY shell session.
- Leave the local project folder in place.
- Keep the OpenAI API key only if you plan to resume the project.
- Keep the $10 hard spend limit enabled while the key remains active.
- Export OPENAI_API_KEY again when you return in a new shell session.
- Restart Streamlit from the activated .venv/ environment when you resume.
- Use the in-app sync button after Streamlit restarts.
Delete - I don't want to use this again
Deleting the project folder is permanent. Use this option only when you no longer need the assistant or its local synchronization history.
- Stop Streamlit by pressing Ctrl+C in the terminal.
- Close the terminal window to end the active OPENAI_API_KEY shell session.
- Delete the project key from the OpenAI API key page.
- Use Finder to locate the folder containing rag_core.py, app.py, and qdrant_data/.
- Move that containing folder to the trash.
- Empty the trash to remove the source files, .venv/, qdrant_data/, and manifest.json permanently.
- Confirm the folder no longer appears in its former Finder location.
Nice Work!
Nice Work!
You did it! You built a self-updating RAG assistant connected to live GitHub Actions documentation.
You've learned how to:
- Connected the GitHub Contents API to three live documentation files. You used each blob SHA as a source version for change detection.
- Built an incremental synchronization pipeline that skips unchanged sources. It replaces changed chunks. It deletes retired vectors. Qdrant stores the synchronized vectors while manifest.json records their versions and point IDs.
- Created a source-aware Streamlit assistant that retrieves evidence from the persistent index. Its answers clearly separate official GitHub Docs from client assumptions.
- Completed the Secret Mission by retiring one configured source. You proved its vectors disappeared. You restored the source with its current GitHub blob SHA. The index converged on the active source list again.
Ready to quiz yourself?