Is Better Search Worth the Wait?

Measure dense, sparse, hybrid, and reranked search against human labels.

Introduction

30 Second Summary

A new search feature can look impressive in a demo while quietly taking longer every time someone uses it. Without a fair comparison, nobody can tell whether the extra work earns its place.

In this project, you will build a local decision review that compares dense search, sparse search, hybrid search, and cross-encoder reranking against human relevance labels. What this does NOT fix: it measures retrieval on public claims, not clinical guidelines, on 20 queries, with no answer generated.

What You'll Build

Your local readout gives you the evidence to decide whether the reranker earns its measured CPU cost on the labelled sample.

By the end of this project, you'll have:

  • A four-row benchmark you can run to compare Recall@5, Recall@50, and nDCG@5 across dense, sparse, hybrid, and hybrid plus rerank.
  • A measured decision trail you can use to weigh signed change over dense against p95 rerank time.
  • A tagged local repository whose self-contained HTML readout opens without a server while keeping every shipped claim traceable.
  • Secret Mission: Score all 300 SciFact test queries after close-out to decide whether the measured lift holds beyond the original 20-query sample.

Are there any prerequisites?

You need Python 3.12 or 3.13, Git, a local editor, and an AI client. Day Zero handles the larger local downloads before the timed session begins.

Before We Start

Before the hands-on work begins, commit to measuring dense search, sparse search, hybrid search, and hybrid-plus-rerank retrieval against human SciFact labels, with unmeasured quality or latency claims treated as owed. What this does NOT fix: it measures retrieval on public claims, not clinical guidelines, on 20 queries, with no answer generated.

Prepare the Local Benchmark

Retrieval scores only mean something when every method runs in the same controlled local environment. A missing tool or incomplete model download would turn the later comparison into an environment test.

Day Zero moves the long package and model downloads ahead of the timed work. The timed session still begins with five minutes of checks that prove the active environment and SciFact data are ready.

In this step, get ready to:
  • Verify Python, Git, Git identity, your editor, and your AI client.
  • Create the local repository, virtual environment, package installation, and model cache.
  • Start the timed session and extract the SciFact benchmark data.
Verify your local prerequisites

Python runs every scorer in this benchmark. Git preserves the design history and build history that make the result auditable.

This project requires Python 3.12 or 3.13. The identity checks also prevent your first commit from stopping to ask who created it.

  • Check your Python version by running the command for your platform:

macOS

python3 --version

What does this command do?

The command prints the version used when you invoke python3. This is the interpreter that creates the project environment on macOS.

Windows

python --version

What does this command do?

The command prints the version used when you invoke python. This is the interpreter that creates the project environment on Windows.

Linux

python3 --version

What does this command do?

The command prints the version used when you invoke python3. This is the interpreter that creates the project environment on Linux.

✔️ I see version 3.12 or 3.13

Your Python version is ready for the pinned package lines. Continue with the Git checks.

ⓧ I see an older version

The benchmark depends on Python 3.12 or 3.13. Install one of those versions before creating the virtual environment.

  • Use the official Python download page to install Python 3.12 or 3.13.
  • Return to your terminal after the installation completes.
  • Run the platform check again. Continue when it reports Python 3.12 or 3.13.

ⓧ The command is unavailable

Python must be installed before this local benchmark can create its isolated environment.

  • Use the official Python download page to install Python 3.12 or 3.13.
  • Start a new terminal window after the installation completes.
  • Run the platform check again. Continue when it reports a supported version.
  • Check Git and both identity values by running:
git --version
git config --get user.name
git config --get user.email

What do these checks prove?

  • The first command confirms that the Git command-line tool is available.
  • The second command prints the name Git records on commits.
  • The third command prints the email Git records on commits.

✔️ Git and both identity values appear

Git is ready to initialize the repository. Both identity values are already available for its future commits.

ⓧ Git is unavailable

Git must be available before you create the local project history.

  • Use the official Git installation page to install Git for your operating system.
  • Start a new terminal window after the installation completes.
  • Run all three checks again.

ⓧ A name or email is blank

A blank line means that identity value is missing. You will set only the missing value inside rerank-day so the change stays local to this project.

  • Use your operating system's app search to confirm your existing editor starts.
  • Use your usual access point to confirm your existing AI client is available for the later close-out.

You now know which local tools are ready. Any missing Git identity value has a project-scoped fix in the next substep.

Prepare the repository and cache

A virtual environment keeps this benchmark's packages separate from your system Python. The local repository gives every later design and build task one traceable home.

  • Move to your Desktop, initialize rerank-day, and enter the repository by running:
cd ~/Desktop
git init rerank-day
cd rerank-day

What do these commands do?

  • The first command places the project in a folder you can find again.
  • The second command creates rerank-day with a new local Git repository.
  • The third command makes rerank-day the current terminal location.

Git reports that it initialized an empty repository. No remote is added, so the benchmark remains local.

  • Set the project-only Git name by running this command if the earlier name check was blank:
git config user.name "Learner"

Why is this setting local?

Running the command inside rerank-day without a global option stores the name for this repository. It leaves your other repositories unchanged.

  • Set the project-only Git email by running this command if the earlier email check was blank:
git config user.email "learner@example.invalid"

Why use this email value?

The value gives this local repository a complete commit identity. It does not connect the repository to a remote account.

  • Confirm both identity values now resolve inside rerank-day by running:
git config --get user.name
git config --get user.email

What should this confirm?

Both commands should now print a value. Git has the identity information required for the commits you create in later steps.

  • Create and activate .venv by running the commands for your platform:

macOS

python3 -m venv .venv
source .venv/bin/activate

What do these commands do?

The first command creates an isolated Python environment inside .venv. The second command updates this shell so its Python and package commands use that environment.

Windows

python -m venv .venv
.\.venv\Scripts\Activate.ps1

What do these commands do?

The first command creates an isolated Python environment inside .venv. The second command activates that environment in the current PowerShell window.

Linux

python3 -m venv .venv
source .venv/bin/activate

What do these commands do?

The first command creates an isolated Python environment inside .venv. The second command updates this shell so its Python and package commands use that environment.

Your terminal prompt should now begin with (.venv). That prefix confirms this terminal uses the local environment.

  • Apply the process-scoped fallback by running these commands only if PowerShell refuses to activate .venv:
Set-ExecutionPolicy -Scope Process -ExecutionPolicy Bypass
.\.venv\Scripts\Activate.ps1

How does the fallback work?

The first command changes the execution policy only for the current PowerShell process. The second command retries activation without changing the permanent machine policy.

Still missing the environment prefix?

  • Confirm your terminal is inside the rerank-day folder before activating the environment.
  • Confirm the environment folder is named exactly .venv.
  • Create the environment again with the platform command if its first creation reported an error.

PyTorch supplies the CPU tensor operations used by the model stack. Sentence Transformers loads the dense model and cross-encoder reranker.

bm25s provides the sparse search stage. The installation also resolves NumPy and Hugging Face Hub CLI dependencies used by the benchmark.

This is the first slow part of Day Zero. Expect several minutes of download output while the CPU package and supporting libraries arrive.

  • Install CPU PyTorch 2.14, Sentence Transformers 6.1, and bm25s 0.3 by running the commands for your platform:

macOS

python3 -m pip install "torch==2.14.*" --index-url https://download.pytorch.org/whl/cpu
python3 -m pip install "sentence-transformers==6.1.*" "bm25s==0.3.*"

What do these commands install?

The first command installs the pinned CPU release line of PyTorch from its CPU package index. The second command installs the pinned Sentence Transformers and bm25s release lines into the active environment.

Windows

python -m pip install "torch==2.14.*" --index-url https://download.pytorch.org/whl/cpu
python -m pip install "sentence-transformers==6.1.*" "bm25s==0.3.*"

What do these commands install?

The first command installs the pinned CPU release line of PyTorch from its CPU package index. The second command installs the pinned Sentence Transformers and bm25s release lines into the active environment.

Linux

python3 -m pip install "torch==2.14.*" --index-url https://download.pytorch.org/whl/cpu
python3 -m pip install "sentence-transformers==6.1.*" "bm25s==0.3.*"

What do these commands install?

The first command installs the pinned CPU release line of PyTorch from its CPU package index. The second command installs the pinned Sentence Transformers and bm25s release lines into the active environment.

The terminal should finish both installations without returning an error. Your active .venv now contains the local benchmark dependencies.

Seeing an installation failure?

  • Confirm (.venv) appears at the start of your terminal prompt.
  • Confirm your Python check reported 3.12 or 3.13.
  • Retry the failed package line after confirming your internet connection is active.

A local model cache keeps the timed run from waiting for model downloads. The dense model and reranker must both finish downloading during Day Zero.

These are the large Day Zero downloads. Let each command finish before starting the next one.

  • Cache both supplied models by running:
hf download BAAI/bge-small-en-v1.5
hf download BAAI/bge-reranker-v2-m3

What is being cached?

  • The BAAI/bge-small-en-v1.5 files support dense semantic retrieval.
  • The BAAI/bge-reranker-v2-m3 files support cross-encoder reranking over the candidate set.

Each download prints progress before returning control to the terminal. Returning without an error confirms the model files reached the local cache.

Did a model download stop early?

  • Confirm the package installation completed before using the Hugging Face Hub command.
  • Retry only the model command that did not finish.
  • Keep the terminal open until the command returns to the prompt without an error.
Complete the timed setup

The timed session starts from a fresh terminal so cached files cannot hide an inactive environment. This five-minute pass proves the repository, package line, and benchmark data are ready together.

  • Use your operating system's app search to start a new terminal window.
  • Enter rerank-day and activate its environment by running the commands for your platform:

macOS

cd ~/Desktop/rerank-day
source .venv/bin/activate

What do these commands restore?

The first command returns to the local repository. The second command makes its isolated Python environment active in this fresh terminal.

Windows

Set-Location "$HOME\Desktop\rerank-day"
.\.venv\Scripts\Activate.ps1

What do these commands restore?

The first command returns to the local repository. The second command makes its isolated Python environment active in this fresh PowerShell window.

Linux

cd ~/Desktop/rerank-day
source .venv/bin/activate

What do these commands restore?

The first command returns to the local repository. The second command makes its isolated Python environment active in this fresh terminal.

  • Inspect the Sentence Transformers package from the active environment by running the command for your platform:

macOS

python3 -m pip show sentence-transformers

What does this inspection prove?

The command asks the active Python environment for the installed package metadata. Its Version line confirms which Sentence Transformers release the benchmark uses.

Windows

python -m pip show sentence-transformers

What does this inspection prove?

The command asks the active Python environment for the installed package metadata. Its Version line confirms which Sentence Transformers release the benchmark uses.

Linux

python3 -m pip show sentence-transformers

What does this inspection prove?

The command asks the active Python environment for the installed package metadata. Its Version line confirms which Sentence Transformers release the benchmark uses.

You should see a Version line beginning with 6.1. That result proves the fresh terminal is using the prepared environment.

The SciFact archive contains the corpus, the queries, and the human qrels used for evaluation. Extracting them under data/scifact gives every retrieval method the same evidence.

The archive download can spend a short while showing progress. Keep the terminal open until extraction returns you to the prompt.

  • Create data, download SciFact, and extract the archive by running the commands for your platform:

macOS

mkdir data
curl --fail -L -o data/scifact.zip https://public.ukp.informatik.tu-darmstadt.de/thakur/BEIR/datasets/scifact.zip
python3 -m zipfile -e data/scifact.zip data

What do these commands create?

  • The first command creates the local data directory.
  • The second command downloads the BEIR SciFact archive into that directory.
  • The third command uses Python's ZIP support to extract the archive under data.

Windows

mkdir data
curl.exe --fail -L -o data/scifact.zip https://public.ukp.informatik.tu-darmstadt.de/thakur/BEIR/datasets/scifact.zip
python -m zipfile -e data/scifact.zip data

What do these commands create?

  • The first command creates the local data directory.
  • The second command downloads the BEIR SciFact archive into that directory.
  • The third command uses Python's ZIP support to extract the archive under data.

Linux

mkdir data
curl --fail -L -o data/scifact.zip https://public.ukp.informatik.tu-darmstadt.de/thakur/BEIR/datasets/scifact.zip
python3 -m zipfile -e data/scifact.zip data

What do these commands create?

  • The first command creates the local data directory.
  • The second command downloads the BEIR SciFact archive into that directory.
  • The third command uses Python's ZIP support to extract the archive under data.
  • Open rerank-day/data/scifact in Finder on macOS, File Explorer on Windows, or your file manager on Linux.
  • Confirm the folder contains corpus.jsonl, queries.jsonl, and qrels/test.tsv.

Before the final inspection, do you expect this fresh terminal to report the pinned Sentence Transformers release line?

  • Run the final package inspection for your platform:

macOS

python3 -m pip show sentence-transformers

What does the final inspection check?

The command reads package metadata from the active environment after the data setup. It confirms the timed session still points at the prepared benchmark environment.

Windows

python -m pip show sentence-transformers

What does the final inspection check?

The command reads package metadata from the active environment after the data setup. It confirms the timed session still points at the prepared benchmark environment.

Linux

python3 -m pip show sentence-transformers

What does the final inspection check?

The command reads package metadata from the active environment after the data setup. It confirms the timed session still points at the prepared benchmark environment.

You should see a Version line beginning with 6.1. That is the final proof that the active environment matches the benchmark's pinned Sentence Transformers release line.

✔️ Awesome, I've got everything!

Your local benchmark is ready. Keep the active terminal available for the design work in the next step.

ⓧ I'd like to double check the full code

Compare your local setup with these required artifact states:

  • .venv: Active local Python virtual environment containing torch==2.14.*, sentence-transformers==6.1.*, bm25s==0.3.*, NumPy, and Hugging Face Hub CLI dependencies.
  • Hugging Face model cache: BAAI/bge-small-en-v1.5 and BAAI/bge-reranker-v2-m3 downloaded locally.
  • data/scifact: Extracted BEIR SciFact dataset, including corpus.jsonl, queries.jsonl, and qrels/test.tsv.
  • rerank-day Git repository: Initialized local Git repository with no remote configured.

That's the slow setup complete: your local environment, model cache, and labelled dataset are ready. Next, you'll freeze the solution contract in five design documents before implementation begins.

Commit the Design Evidence

Your local benchmark environment is ready. Before you write retrieval code, you need a record of what you expect the benchmark to prove.

A local Git commit freezes the solution contract before results can influence the reasoning. It preserves the alternatives, costs, reversal triggers, topology, timing, sources, and limits behind the comparison.

In this step, get ready to:
  • Build a project file set that explains the review and excludes local artifacts.
  • Preserve five supplied design documents without revising their evidence.
  • Create a Git baseline that proves the design existed before implementation.
Create the root project files

The root files explain the decision review at a glance. They also prevent local environments, downloaded data, cached embeddings, and generated Python files from entering the repository.

  • Press Cmd+Space (macOS) or the Windows key (Windows) to open your operating system's search bar.
  • Type the name of your existing editor into the search bar.
  • Press Enter to open the editor.
  • Use the editor's folder picker to select the rerank-day folder on your Desktop.
  • Create README.md inside rerank-day.
  • Add sections named What, Why, and How.
  • Write one short paragraph under each heading to cover the four-row review, the quality-versus-CPU decision, and the evidence-first build order.
  • Save README.md.

You should see What, Why, and How in README.md.

  • Create TASKS.md inside rerank-day by copying this ordered task list:
1. Prepare Day Zero and complete the five-minute timed setup.
2. Copy and commit the five design documents.
3. Build and commit the scorer and label loader.
4. Build and commit dense, sparse, hybrid, and reranking stages, then preserve the must-fail evidence.
5. Restore the scorer, evaluate, score and clean the three close-out drafts, then tag the local release.
6. After close-out only, optionally score all 300 test queries and decide whether the lift holds.

What does this task list protect?

The order keeps design evidence ahead of implementation. It also keeps the optional 300-query evaluation behind the tagged 20-query close-out.

  • Save TASKS.md.
  • Confirm that TASKS.md contains six numbered tasks.

You should see setup first and the optional 300-query evaluation last.

  • Create .gitignore inside rerank-day with this content:
.venv/
data/
corpus_emb.npy
__pycache__/

What does this file exclude?

  • .venv/ keeps the local virtual environment outside version control.
  • data/ keeps the downloaded SciFact files outside version control.
  • corpus_emb.npy keeps the generated dense embedding cache outside version control.
  • __pycache__/ keeps generated Python cache folders outside version control.
  • Save .gitignore.
  • Confirm that your editor lists README.md, TASKS.md, and .gitignore at the top level of rerank-day.

Why freeze design first?

A pre-implementation design records what you expected before the benchmark produced results. That ordering makes later explanations easier to audit.

The documents also name the evidence limit up front. Twenty public-claim queries cannot establish a general or clinical win.

Copy the five design documents

These five supplied documents form the design contract. Copy them exactly so implementation cannot quietly change the benchmark question or its acceptance criteria.

  • Create the docs folder inside rerank-day by running this command:
mkdir docs

You should see an empty docs folder in your editor's file tree.

  • Create docs/requirements-brief.md with the following supplied content:
# The reranker that has to pay for itself

## Problem

Determine whether cross-encoder reranking improves top-five retrieval quality enough to justify its measured CPU time.

## Build

Compare dense, sparse, hybrid, and hybrid plus rerank over the same SciFact corpus, first 20 test queries, human qrels, and top-50 candidate depth.

## Acceptance

- The scorer logs three passing fixtures.
- The intended nDCG mutation logs got 0.6509 against expected 0.4982.
- The scorer is restored and logs all checks passing again.
- Evaluation logs four rows, Recall@5, Recall@50, nDCG@5, signed changes from dense, same-50 equality, and p95 rerank time.
- Close-out drafts a teach-back, after-action review, and self-contained HTML readout, then removes or strikes every untraceable claim.

## Limits

What this does NOT fix: it measures retrieval on public claims, not clinical guidelines, on 20 queries, with no answer generated.

What does the requirements brief fix in place?

The brief defines the quality-versus-CPU problem. It also fixes the four search rows, the human relevance labels, the candidate depth, the acceptance evidence, and the limitation sentence.

  • Save docs/requirements-brief.md.
  • Confirm that the file ends with the sentence beginning What this does NOT fix.

You should see the limitation name public claims, clinical guidelines, 20 queries, and the absence of answer generation.

  • Create docs/decisions.md with the following supplied content:
# Decision Records

## Record 1

Choice: Score dense alone first on human labels. After the run, add its measured Recall@5, Recall@50, and nDCG@5 here.

Alternative beaten: Report only the hybrid result.

Cost: A noisy 20-query sample.

Reversal trigger: Replace it with the product's own labelled questions.

## Record 2

Choice: Use RRF with constant 60 over dense and sparse top-50 lists, then let the cross-encoder reorder those same 50 rows.

Alternative beaten: Sum raw scores that live on different scales.

Cost: CPU seconds per query.

Reversal trigger: Reverse if score-weighted fusion tuned on different queries beats RRF, or if p95 rerank time breaks the product budget.

## Record 3

Choice: Use one library, bm25s, for exact sparse search.

Alternative beaten: A vector database plus Ollama.

Cost: Index tuning and a generated answer are left out.

Reversal trigger: Reverse when the corpus outgrows exact search.

What do the decision records capture?

Each record connects a choice to an alternative, cost, and reversal trigger. The choices cover the dense search baseline, Reciprocal Rank Fusion over dense and sparse search results, same-candidate cross-encoder reranking, and exact search with bm25s.

  • Save docs/decisions.md.
  • Confirm that the file contains three numbered decision records.

You should see four fields in every record: choice, alternative beaten, cost, and reversal trigger.

  • Create docs/topology.md with the following supplied content:
# Topology

```mermaid
flowchart LR
    C[SciFact corpus] --> L[Load 5,183 documents]
    Q[First 20 test queries] --> L
    H[Human qrels] --> L
    L --> D[Dense top 50]
    L --> S[Sparse top 50]
    D --> F[RRF constant 60]
    S --> F
    F --> Y[Hybrid top 50]
    Y --> X[Cross-encoder reorders same 50]
    D --> E[Four-row evaluation]
    S --> E
    Y --> E
    X --> E
    E --> J[results.json]
    E --> C0[Close-out drafts and trace review]
```

What does the topology show?

The Mermaid flowchart maps one labelled dataset into dense and sparse top-50 lists. It then shows rank fusion, same-50 reranking, four-row evaluation, measured results, and the claim review.

  • Save docs/topology.md.
  • Confirm that the final two nodes point to results.json and the close-out trace review.

You should see the cross-encoder receive only the hybrid top 50. That boundary prevents reranking from adding a new candidate.

  • Create docs/build-plan.md with the following supplied content:
# Build Plan

| Phase | Minutes | Exit evidence |
| --- | ---: | --- |
| Setup | 5 | Environment active, Sentence Transformers 6.1 shown, SciFact extracted |
| Design | 7 | Five supplied documents copied and committed before code |
| Build | 20 | Scorer, loader, dense, sparse, hybrid, reranker, and evaluator committed by task |
| Validation | 8 | PASS, intended FAIL, restoration, and four-row result preserved |
| Close-out | 5 | Three drafts traced, cleaned, committed, and tagged |
| Total | 45 | One local solution review complete |

The optional 300-query Secret Mission starts only after close-out.

What does the build plan control?

The table allocates the complete 45-minute session across setup, design, build, validation, and close-out. It keeps the optional 300-query run outside the required release.

  • Save docs/build-plan.md.
  • Confirm that the Total row shows 45 minutes.

You should see the optional Secret Mission begin only after close-out.

  • Create docs/source-map.md with the following supplied content:
# Source Map

## Models and data

- https://huggingface.co/BAAI/bge-small-en-v1.5
- https://huggingface.co/BAAI/bge-reranker-v2-m3
- https://huggingface.co/datasets/BeIR/scifact
- https://download.pytorch.org/whl/cpu/torch/

## Dated implementation reason

Sentence Transformers v6.0.0 was published on 18 August 2026. CrossEncoder.rank score values became Python floats that are directly JSON serializable.

- https://github.com/huggingface/sentence-transformers/releases/tag/v6.0.0

## Readout context figure

WHO reported a projected global shortage of 11.1 million health workers by 2030 on 18 September 2026. This is context only and does not turn SciFact into clinical-guideline evidence.

- https://www.who.int/news/item/18-09-2026-one-in-four-doctors-nearing-retirement-age-as-who-warns-of-global-health-workforce-gap

What does the source map separate?

The source map records the models, SciFact data, and CPU package source. It also preserves the dated Sentence Transformers v6.0.0 implementation reason.

The WHO figure remains context for the final readout. The document explicitly prevents that figure from turning SciFact into clinical-guideline evidence.

  • Save docs/source-map.md.
  • Confirm that the file contains sections for models and data, the dated implementation reason, and the readout context figure.

You should now see all five supplied documents inside docs.

✔️ Awesome, I've got everything!

Your root files and five supplied design documents are ready. Make sure every file is saved before you create the first commit.

  • Confirm that README.md contains What, Why, and How sections.
  • Confirm that TASKS.md contains the six ordered project tasks.
  • Confirm that .gitignore contains four ignore rules.
  • Confirm that the docs folder contains exactly five design documents.

ⓧ I'd like to double check the full code

Compare your files with these complete references. Keep your own project-specific wording under the What, Why, and How sections in README.md.

1. Prepare Day Zero and complete the five-minute timed setup.
2. Copy and commit the five design documents.
3. Build and commit the scorer and label loader.
4. Build and commit dense, sparse, hybrid, and reranking stages, then preserve the must-fail evidence.
5. Restore the scorer, evaluate, score and clean the three close-out drafts, then tag the local release.
6. After close-out only, optionally score all 300 test queries and decide whether the lift holds.

This reference keeps the required work in order and places the optional 300-query run after close-out.

.venv/
data/
corpus_emb.npy
__pycache__/

This reference excludes the active environment, downloaded data, generated embeddings, and Python cache folders.

# The reranker that has to pay for itself

## Problem

Determine whether cross-encoder reranking improves top-five retrieval quality enough to justify its measured CPU time.

## Build

Compare dense, sparse, hybrid, and hybrid plus rerank over the same SciFact corpus, first 20 test queries, human qrels, and top-50 candidate depth.

## Acceptance

- The scorer logs three passing fixtures.
- The intended nDCG mutation logs got 0.6509 against expected 0.4982.
- The scorer is restored and logs all checks passing again.
- Evaluation logs four rows, Recall@5, Recall@50, nDCG@5, signed changes from dense, same-50 equality, and p95 rerank time.
- Close-out drafts a teach-back, after-action review, and self-contained HTML readout, then removes or strikes every untraceable claim.

## Limits

What this does NOT fix: it measures retrieval on public claims, not clinical guidelines, on 20 queries, with no answer generated.

This reference fixes the benchmark problem, build, acceptance checks, and evidence limit.

# Decision Records

## Record 1

Choice: Score dense alone first on human labels. After the run, add its measured Recall@5, Recall@50, and nDCG@5 here.

Alternative beaten: Report only the hybrid result.

Cost: A noisy 20-query sample.

Reversal trigger: Replace it with the product's own labelled questions.

## Record 2

Choice: Use RRF with constant 60 over dense and sparse top-50 lists, then let the cross-encoder reorder those same 50 rows.

Alternative beaten: Sum raw scores that live on different scales.

Cost: CPU seconds per query.

Reversal trigger: Reverse if score-weighted fusion tuned on different queries beats RRF, or if p95 rerank time breaks the product budget.

## Record 3

Choice: Use one library, bm25s, for exact sparse search.

Alternative beaten: A vector database plus Ollama.

Cost: Index tuning and a generated answer are left out.

Reversal trigger: Reverse when the corpus outgrows exact search.

This reference preserves all three choices with their alternatives, costs, and reversal triggers.

# Topology

```mermaid
flowchart LR
    C[SciFact corpus] --> L[Load 5,183 documents]
    Q[First 20 test queries] --> L
    H[Human qrels] --> L
    L --> D[Dense top 50]
    L --> S[Sparse top 50]
    D --> F[RRF constant 60]
    S --> F
    F --> Y[Hybrid top 50]
    Y --> X[Cross-encoder reorders same 50]
    D --> E[Four-row evaluation]
    S --> E
    Y --> E
    X --> E
    E --> J[results.json]
    E --> C0[Close-out drafts and trace review]
```

This reference maps the shared data, four search rows, measured result file, and close-out review.

# Build Plan

| Phase | Minutes | Exit evidence |
| --- | ---: | --- |
| Setup | 5 | Environment active, Sentence Transformers 6.1 shown, SciFact extracted |
| Design | 7 | Five supplied documents copied and committed before code |
| Build | 20 | Scorer, loader, dense, sparse, hybrid, reranker, and evaluator committed by task |
| Validation | 8 | PASS, intended FAIL, restoration, and four-row result preserved |
| Close-out | 5 | Three drafts traced, cleaned, committed, and tagged |
| Total | 45 | One local solution review complete |

The optional 300-query Secret Mission starts only after close-out.

This reference preserves the complete 45-minute plan and the post-close-out boundary.

# Source Map

## Models and data

- https://huggingface.co/BAAI/bge-small-en-v1.5
- https://huggingface.co/BAAI/bge-reranker-v2-m3
- https://huggingface.co/datasets/BeIR/scifact
- https://download.pytorch.org/whl/cpu/torch/

## Dated implementation reason

Sentence Transformers v6.0.0 was published on 18 August 2026. CrossEncoder.rank score values became Python floats that are directly JSON serializable.

- https://github.com/huggingface/sentence-transformers/releases/tag/v6.0.0

## Readout context figure

WHO reported a projected global shortage of 11.1 million health workers by 2030 on 18 September 2026. This is context only and does not turn SciFact into clinical-guideline evidence.

- https://www.who.int/news/item/18-09-2026-one-in-four-doctors-nearing-retirement-age-as-who-warns-of-global-health-workforce-gap

This reference preserves the model, data, package, release, and readout-context sources.

Missing a document or heading?

  • Check that the folder name is exactly docs in lowercase.
  • Check each filename against the full-code reference without adding spaces or changing hyphens.
  • Compare the last line of each document because copied blocks are easiest to truncate at the bottom.

Still stuck? Help me compare my rerank-day design files with the supplied references.

Commit and defend the design

A commit turns the supplied design into a fixed point in local history. The retrieval code added later can now be compared against decisions that already existed.

  • Stage every root file and design document by running:
git add -A

What does this command do?

git add -A stages additions, modifications, and removals. The ignore rules keep .venv, data, corpus_emb.npy, and __pycache__ outside the commit.

  • Confirm exactly which files Git staged by running:
git status --short

You should see README.md, TASKS.md, .gitignore, and five files under docs. You should not see .venv, data, corpus_emb.npy, or __pycache__.

  • Create the first local commit by running:
git commit -m "Add project docs"

What does this command record?

The command records the staged design evidence with the message Add project docs. This commit becomes the pre-implementation baseline.

Before you check the history, what message do you expect the latest entry to contain?

  • Inspect the latest history entry by running:
git log -1 --oneline

What does this command show?

The command prints the latest commit as a shortened identifier followed by its message. It gives you visible proof that design was committed before implementation.

You should see one line ending with Add project docs. That line confirms your first commit preserves the design evidence.

Did the commit fail?

  • Check that your terminal is still inside the rerank-day repository.
  • Check that every file is saved in your editor before staging again.
  • Check that your local Git name and email still have values if the commit command reports an identity problem.

Need a hand? Help me troubleshoot why my first rerank-day Git commit did not complete.

The commit proves when the design existed. The final check is whether you can defend the reasoning inside it without relying on later benchmark results.

  • Read Decision Record 1 aloud using its choice, alternative beaten, cost, and reversal trigger.
  • Read Decision Record 2 aloud using its choice, alternative beaten, cost, and reversal trigger.
  • Read Decision Record 3 aloud using its choice, alternative beaten, cost, and reversal trigger.

That is the evidence boundary secured. Your repository now has a reviewable design baseline that predates every implementation choice.

Your design evidence is fixed in local history. Next, you will make the scorer and human-label loader executable before any search method can make a quality claim.

Build the Scorer and Loader

Your design evidence is committed in Git. The first implementation claim now needs an executable measuring stick.

Human relevance labels provide the reference for that measurement. The scorer must work before any search method can claim a quality improvement.

A shared SciFact loader keeps the comparison consistent. Every later method receives the same 5,183 documents and the same 20 labelled queries.

In this step, get ready to:
  • Implement Recall@5, nDCG@5, and Reciprocal Rank Fusion.
  • Prove the scorer with three passing fixtures.
  • Commit a loader for the first 20 human-labelled test queries.
Define the scoring functions

A benchmark needs metric definitions that behave identically for every search method. The first two functions measure coverage and top-five ranking quality.

  • Return to the rerank-day folder in your editor.
  • Create metrics.py inside rerank-day using your editor's new-file control.
  • Add the imports and metric functions by copying this code into metrics.py:
import math
import sys


def recall_at_k(ranked, relevant, k):
    # An empty label set has no relevant documents to recover.
    if not relevant:
        return 0.0

    # Count unique relevant documents within the first k results.
    retrieved = set(ranked[:k]) & set(relevant)
    return len(retrieved) / len(set(relevant))

What does Recall@k measure?

  • The ranked list holds document identifiers in result order.
  • The relevant set holds the human-labelled documents for the query.
  • The function divides the relevant documents found in the first k positions by the total number of relevant documents.
  • Save metrics.py.
  • Confirm your editor shows recall_at_k() in the file outline with no syntax errors.

Is Recall@k flagged?

  • Check that the function body uses four spaces of indentation.
  • Check that the set intersection uses one & operator.
  • Check that retrieved is defined before the final return statement.

Still stuck? Help me debug the Recall@k function.

Recall measures coverage, but it treats every position inside the cutoff equally. ndcg_at_k() adds ranking quality by rewarding relevant documents that appear earlier.

  • Add the ranking-quality function below recall_at_k() by copying this code:
def ndcg_at_k(ranked, relevant, k):
    # Use a set for fast membership checks while scoring the ranking.
    relevant = set(relevant)
    gain = sum(
        1 / math.log2(rank + 1)
        for rank, doc_id in enumerate(ranked[:k], start=1)
        if doc_id in relevant
    )

    # Compare the measured gain with the best possible order.
    ideal_hits = min(k, len(relevant))
    ideal = sum(1 / math.log2(rank + 1) for rank in range(1, ideal_hits + 1))
    return 0.0 if ideal == 0 else gain / ideal

How does nDCG@k score ranking quality?

  • The gain value discounts relevant documents as their rank moves lower.
  • The ideal value represents the best possible order at the same cutoff.
  • Dividing the measured gain by the ideal gain produces a score between zero and one.
  • Save metrics.py.
  • Confirm your editor shows ndcg_at_k() in the file outline with no syntax errors.

Is nDCG@k flagged?

  • Check that enumerate() starts at 1.
  • Check that both sum() expressions close before the final return statement.
  • Check that ideal_hits uses the smaller of k and the number of relevant documents.

Need another pair of eyes? Help me debug the nDCG@k function.

Dense and sparse scores use different scales. rrf() combines their rank positions with the verified constant 60.

  • Add rank fusion below ndcg_at_k() by copying this code:
def rrf(lists, k=60):
    # Accumulate one reciprocal-rank contribution per input list.
    scores = {}
    first_seen = {}
    seen_count = 0
    for ranked in lists:
        for rank, item in enumerate(ranked, start=1):
            if item not in first_seen:
                first_seen[item] = seen_count
                seen_count += 1
            scores[item] = scores.get(item, 0.0) + 1 / (k + rank)

    # Use first appearance to make tied results deterministic.
    return sorted(scores, key=lambda item: (-scores[item], first_seen[item]))

How does rank fusion work?

  • The scores dictionary adds each item's contribution across the supplied rankings.
  • The first_seen dictionary records a stable order for tied scores.
  • The final sort places the highest fused score first.
  • Save metrics.py.
  • Confirm your editor shows rrf() in the file outline with no syntax errors.

Is rank fusion flagged?

  • Check that both loops remain inside rrf().
  • Check that the reciprocal contribution uses k + rank.
  • Check that the final sort uses negative score before first_seen.

Still stuck? Help me debug the rank fusion function.

The fixture runner needs one consistent way to compare measured values with expectations. check() prints a clear result and returns whether the fixture passed.

  • Add the fixture comparison helper below rrf() by copying this code:
def check(name, got, expected):
    # Compare lists exactly and numeric values to four decimal places.
    passed = got == expected if isinstance(expected, list) else round(got, 4) == round(expected, 4)
    print(f"{'PASS' if passed else 'FAIL'} {name}: got {got}, expected {expected}")
    return passed

What does the fixture helper check?

  • The passed value records whether the measured and expected results agree.
  • List results require an exact order match.
  • Numeric results are compared at four decimal places to match the fixture output.
  • Save metrics.py.
  • Confirm your editor shows check() in the file outline with no syntax errors.

Is the fixture helper flagged?

  • Check that the list comparison uses exact equality.
  • Check that both numeric values pass through round() with 4.
  • Check that the function returns passed after printing the result.

Need help? Help me debug the fixture comparison helper.

Run the fixtures and commit the scorer

Small fixtures give each function a known input with a known result. A later defect becomes visible when one of these expectations fails.

  • Add the fixture runner by copying this code at the bottom of metrics.py:
if __name__ == "__main__":
    # Use fixed inputs so every run tests the same metric behavior.
    ranked = ["d3", "d1", "d7", "d2", "d9"]
    relevant = {"d1", "d2", "d4"}
    checks = [
        check("recall@5", recall_at_k(ranked, relevant, 5), 0.6667),
        check("nDCG@5", ndcg_at_k(ranked, relevant, 5), 0.4982),
        check("rrf", rrf([["x", "y", "z"], ["y", "z", "x"]]), ["y", "x", "z"]),
    ]

    # Return a failing process status when any fixture disagrees.
    if all(checks):
        print("ALL CHECKS PASS")
    else:
        sys.exit(1)

What does the fixture runner prove?

  • The ranked list gives the metrics a fixed result order.
  • The relevant set represents the known labels for the example.
  • The script exits unsuccessfully when any fixture fails.
  • Save metrics.py.

Before you run the scorer, do you expect all three fixtures to agree with their recorded values?

  • Run the scorer for your platform to append its evidence to check-out.txt:

macOS

python3 metrics.py 2>&1 | tee -a check-out.txt

What does this command do?

This runs the scorer with Python. It appends the displayed fixture output to check-out.txt as persistent evidence.

Windows

python metrics.py 2>&1 | Tee-Object -Append check-out.txt

What does this command do?

This runs the scorer with Python. It appends the displayed fixture output to check-out.txt as persistent evidence.

Linux

python3 metrics.py 2>&1 | tee -a check-out.txt

What does this command do?

This runs the scorer with Python. It appends the displayed fixture output to check-out.txt as persistent evidence.

You should see three lines beginning with PASS. The final line should read ALL CHECKS PASS.

Strong start. Your retrieval review now has an executable guard for every quality calculation that follows.

Do any fixtures fail?

  • Match the failed fixture name to its function in metrics.py.
  • Compare its printed value with the matching expected value.
  • Confirm check-out.txt contains the latest run.

Need help interpreting the difference? Help me trace the failed fixture.

  • Stage the scorer evidence by running this command:
git add -A

What does this command do?

This stages the scorer and its recorded output. The next commit preserves them as one implementation task.

  • Confirm the scorer files are staged by running:
git status --short

You should see metrics.py and check-out.txt in the staged changes.

  • Commit the scorer by running this command:
git commit -m "Add the scorer"

What does this command do?

This records the staged snapshot with the message Add the scorer.

You should see a commit summary for Add the scorer.

✔️ Awesome, I've got everything!

Your scorer contains all four functions. Its passing fixture output is preserved in check-out.txt.

ⓧ I'd like to double check the full code

Compare your committed metrics.py with this exact reference.

import math
import sys


def recall_at_k(ranked, relevant, k):
    # An empty label set has no relevant documents to recover.
    if not relevant:
        return 0.0

    # Count unique relevant documents within the first k results.
    retrieved = set(ranked[:k]) & set(relevant)
    return len(retrieved) / len(set(relevant))


def ndcg_at_k(ranked, relevant, k):
    # Use a set for fast membership checks while scoring the ranking.
    relevant = set(relevant)
    gain = sum(
        1 / math.log2(rank + 1)
        for rank, doc_id in enumerate(ranked[:k], start=1)
        if doc_id in relevant
    )

    # Compare the measured gain with the best possible order.
    ideal_hits = min(k, len(relevant))
    ideal = sum(1 / math.log2(rank + 1) for rank in range(1, ideal_hits + 1))
    return 0.0 if ideal == 0 else gain / ideal


def rrf(lists, k=60):
    # Accumulate one reciprocal-rank contribution per input list.
    scores = {}
    first_seen = {}
    seen_count = 0
    for ranked in lists:
        for rank, item in enumerate(ranked, start=1):
            if item not in first_seen:
                first_seen[item] = seen_count
                seen_count += 1
            scores[item] = scores.get(item, 0.0) + 1 / (k + rank)

    # Use first appearance to make tied results deterministic.
    return sorted(scores, key=lambda item: (-scores[item], first_seen[item]))


def check(name, got, expected):
    # Compare lists exactly and numeric values to four decimal places.
    passed = got == expected if isinstance(expected, list) else round(got, 4) == round(expected, 4)
    print(f"{'PASS' if passed else 'FAIL'} {name}: got {got}, expected {expected}")
    return passed


if __name__ == "__main__":
    # Use fixed inputs so every run tests the same metric behavior.
    ranked = ["d3", "d1", "d7", "d2", "d9"]
    relevant = {"d1", "d2", "d4"}
    checks = [
        check("recall@5", recall_at_k(ranked, relevant, 5), 0.6667),
        check("nDCG@5", ndcg_at_k(ranked, relevant, 5), 0.4982),
        check("rrf", rrf([["x", "y", "z"], ["y", "z", "x"]]), ["y", "x", "z"]),
    ]

    # Return a failing process status when any fixture disagrees.
    if all(checks):
        print("ALL CHECKS PASS")
    else:
        sys.exit(1)

The reference preserves the exact function names and fixture values required by later steps.

Load the labelled SciFact data

Every search method needs identical documents and queries. The loader reads the corpus, query text, and qrels from the extracted dataset.

The source files use JSON or tab-separated rows. Converting each identifier to a string keeps comparisons consistent.

  • Create retrieve.py inside rerank-day using your editor's new-file control.
  • Add the imports, data path, and corpus helper by copying this code into retrieve.py:
import csv
import json
from pathlib import Path

# Keep every dataset path relative to the rerank-day folder.
DATA = Path("data/scifact")


def load_corpus():
    # Preserve document IDs beside the text used for retrieval.
    ids = []
    texts = []
    with (DATA / "corpus.jsonl").open(encoding="utf-8") as handle:
        for line in handle:
            row = json.loads(line)
            ids.append(str(row["_id"]))
            texts.append(f"{row.get('title', '')} {row.get('text', '')}".strip())
    return ids, texts

What does the corpus helper return?

  • The DATA constant points to data/scifact.
  • The ids list preserves every document identifier as a string.
  • The texts list joins each title with its document text.
  • Returning both lists keeps document identifiers aligned with retrieval text.
  • Save retrieve.py.
  • Confirm your editor shows load_corpus() in the file outline with no syntax errors.

Is the corpus helper flagged?

  • Check that DATA uses the exact path data/scifact.
  • Check the quotation marks inside the formatted document string.
  • Check that return ids, texts remains inside load_corpus().

Still seeing an editor warning? Help me debug the corpus helper.

The corpus contains the evidence candidates, while the query file contains the questions used to search them. A separate helper keeps query identifiers aligned with their text.

  • Add the query helper below load_corpus() by copying this code:
def load_queries():
    # Map every query ID to the text sent into each search method.
    query_text = {}
    with (DATA / "queries.jsonl").open(encoding="utf-8") as handle:
        for line in handle:
            row = json.loads(line)
            query_text[str(row["_id"])] = row["text"]
    return query_text

What does the query helper return?

  • The helper reads every row from queries.jsonl.
  • Each identifier becomes a string before it enters query_text.
  • The returned dictionary lets the label loader select query text by identifier.
  • Save retrieve.py.
  • Confirm your editor shows load_queries() in the file outline with no syntax errors.

The qrels file connects each query to its human-labelled relevant documents. Preserving first-seen query order creates a repeatable sample.

  • Add the relevance-label helper below load_queries() by copying this code:
def load_labels():
    # Group every positively labelled document under its query ID.
    all_labels = {}
    with (DATA / "qrels" / "test.tsv").open(encoding="utf-8") as handle:
        for row in csv.DictReader(handle, delimiter="\t"):
            qid = str(row["query-id"])
            corpus_id = str(row["corpus-id"])
            all_labels.setdefault(qid, set())
            if int(row["score"]) > 0:
                all_labels[qid].add(corpus_id)
    return all_labels

How are the labels grouped?

  • The helper reads the tab-separated rows from qrels/test.tsv.
  • Each positive score adds a relevant document identifier to its query's set.
  • The dictionary keeps the source file's first-seen query order.
  • Save retrieve.py.
  • Confirm your editor shows load_labels() in the file outline with no syntax errors.

Is the label helper flagged?

  • Check that the delimiter is exactly \t.
  • Check that both identifiers are converted to strings.
  • Check that only scores above zero enter the label set.

Still stuck? Help me debug the label helper.

The three helpers now load each source independently. load() combines them into the shared benchmark interface used by later search stages.

  • Add the shared loader below load_labels() by copying this code:
def load():
    # Assemble the same corpus and labelled sample for every search method.
    ids, texts = load_corpus()
    query_text = load_queries()
    all_labels = load_labels()
    qids = list(all_labels)[:20]
    qtexts = [query_text[qid] for qid in qids]
    labels = {qid: all_labels[qid] for qid in qids}
    return ids, texts, qids, qtexts, labels

What does the shared loader guarantee?

  • The qids slice keeps the first 20 labelled query identifiers.
  • The qtexts list follows the same identifier order.
  • The returned values keep documents, queries, and labels aligned for every method.
  • Save retrieve.py.
  • Confirm your editor shows load() in the file outline with no syntax errors.

Is the shared loader flagged?

  • Check that qids is sliced to 20.
  • Check that qtexts and labels both use those same identifiers.
  • Check that the return statement keeps all five values in the shown order.

Need help? Help me debug the shared loader.

A small main block makes the loader's contract visible from the terminal. It prints the document and query counts without starting a search.

  • Add the count check at the bottom of retrieve.py by copying this code:
if __name__ == "__main__":
    # Print the two counts that define the required benchmark input.
    ids, texts, qids, qtexts, labels = load()
    print(f"{len(ids)} documents")
    print(f"{len(qids)} queries")

What does the count check prove?

  • Calling load() exercises all three source readers.
  • The two printed counts provide a visible check against the design contract.
  • Save retrieve.py.

Before you run the loader, which two counts would prove that its output matches the design contract?

  • Run the loader for your platform to append its counts to check-out.txt:

macOS

python3 retrieve.py 2>&1 | tee -a check-out.txt

What does this command do?

This runs the loader with Python. It appends the document and query counts to the existing scorer evidence.

Windows

python retrieve.py 2>&1 | Tee-Object -Append check-out.txt

What does this command do?

This runs the loader with Python. It appends the document and query counts to the existing scorer evidence.

Linux

python3 retrieve.py 2>&1 | tee -a check-out.txt

What does this command do?

This runs the loader with Python. It appends the document and query counts to the existing scorer evidence.

You should see 5183 documents followed by 20 queries.

That checkpoint locks the benchmark inputs. Every later search method now receives the same corpus and labelled sample.

Do the counts differ?

  • Confirm the extracted dataset remains inside data/scifact.
  • Check that the corpus reader opens corpus.jsonl.
  • Check that the qrels reader opens qrels/test.tsv.

Need help tracing the mismatch? Help me debug the loader counts.

  • Stage the loader evidence by running this command:
git add -A

What does this command do?

This stages retrieve.py and the updated check-out.txt.

  • Confirm the loader files are staged by running:
git status --short

You should see retrieve.py and the updated check-out.txt in the staged changes.

  • Commit the loader by running this command:
git commit -m "Add the SciFact loader"

What does this command do?

This records the staged snapshot with the message Add the SciFact loader.

You should see a commit summary for Add the SciFact loader.

✔️ Awesome, I've got everything!

Your committed loader returns the corpus data and the 20-query benchmark inputs. The recorded counts sit after the scorer output in check-out.txt.

ⓧ I'd like to double check the full code

Compare your committed retrieve.py with this exact reference.

import csv
import json
from pathlib import Path

# Keep every dataset path relative to the rerank-day folder.
DATA = Path("data/scifact")


def load_corpus():
    # Preserve document IDs beside the text used for retrieval.
    ids = []
    texts = []
    with (DATA / "corpus.jsonl").open(encoding="utf-8") as handle:
        for line in handle:
            row = json.loads(line)
            ids.append(str(row["_id"]))
            texts.append(f"{row.get('title', '')} {row.get('text', '')}".strip())
    return ids, texts


def load_queries():
    # Map every query ID to the text sent into each search method.
    query_text = {}
    with (DATA / "queries.jsonl").open(encoding="utf-8") as handle:
        for line in handle:
            row = json.loads(line)
            query_text[str(row["_id"])] = row["text"]
    return query_text


def load_labels():
    # Group every positively labelled document under its query ID.
    all_labels = {}
    with (DATA / "qrels" / "test.tsv").open(encoding="utf-8") as handle:
        for row in csv.DictReader(handle, delimiter="\t"):
            qid = str(row["query-id"])
            corpus_id = str(row["corpus-id"])
            all_labels.setdefault(qid, set())
            if int(row["score"]) > 0:
                all_labels[qid].add(corpus_id)
    return all_labels


def load():
    # Assemble the same corpus and labelled sample for every search method.
    ids, texts = load_corpus()
    query_text = load_queries()
    all_labels = load_labels()
    qids = list(all_labels)[:20]
    qtexts = [query_text[qid] for qid in qids]
    labels = {qid: all_labels[qid] for qid in qids}
    return ids, texts, qids, qtexts, labels


if __name__ == "__main__":
    # Print the two counts that define the required benchmark input.
    ids, texts, qids, qtexts, labels = load()
    print(f"{len(ids)} documents")
    print(f"{len(qids)} queries")

The reference matches the loader state required before retrieval is added.

Your scorer passes all three fixtures. Your loader now supplies 5,183 documents and the same 20 labelled queries to every method.

Next up, these committed foundations become dense, sparse, hybrid, and reranked search stages.

Build and Test the Search Stages

Your scorer and SciFact loader now turn the benchmark contract into executable inputs. The committed design evidence gives every search decision a testable purpose.

This step compares retrieval methods against the same human relevance labels. It also preserves evidence that the scorer catches a faulty metric definition.

In this step, get ready to:
  • Build the dense baseline over the first 20 labelled queries.
  • Add sparse retrieval with rank-based hybrid fusion.
  • Rerank the same 50 candidates before testing the metric guard.
Build the dense baseline

Dense search represents every corpus document as a numeric embedding. Each query retrieves the 50 corpus rows with the highest normalized similarity.

  • Open retrieve.py in your editor.
  • Find the import group and DATA constant shown here:
import csv
import json
from pathlib import Path

# Keep every dataset path relative to the rerank-day folder.
DATA = Path("data/scifact")
  • Replace that block with the dense-search imports and model configuration shown here:
import csv
import json
from pathlib import Path

# Load array operations and the local embedding model.
import numpy as np
from sentence_transformers import SentenceTransformer

# Keep data and model references stable across every benchmark run.
DATA = Path("data/scifact")
DENSE_MODEL = "BAAI/bge-small-en-v1.5"
  • Add dense() below load() by copying this code:
def dense(texts, qtexts):
    # Run every embedding calculation on the learner's CPU.
    model = SentenceTransformer(DENSE_MODEL, device="cpu")
    cache = Path("corpus_emb.npy")

    # Reuse saved corpus embeddings after the first run.
    if cache.exists():
        doc_emb = np.load(cache)
    else:
        doc_emb = model.encode(texts, normalize_embeddings=True, show_progress_bar=True)
        np.save(cache, doc_emb)

    # Rank corpus rows by dot product against normalized query embeddings.
    query_emb = model.encode(qtexts, normalize_embeddings=True, show_progress_bar=False)
    return [np.argsort(-(doc_emb @ query))[:50].tolist() for query in query_emb]

What does the dense function do?

  • NumPy stores the cached embeddings. It also calculates the dot products used for ranking.
  • Sentence Transformers loads BAAI/bge-small-en-v1.5 on the CPU.
  • normalize_embeddings=True gives every vector unit length. The dot product can then represent cosine similarity.
  • corpus_emb.npy prevents later runs from encoding all 5,183 documents again.
  • Save retrieve.py.
  • Confirm your editor lists dense() in the file outline without a syntax warning.

Is the dense function flagged?

  • Check that DENSE_MODEL appears above the function definitions.
  • Check that both model.encode() calls use normalize_embeddings=True.
  • Check that the return statement keeps the first 50 row indices.

Need help comparing the function? Help me debug the dense function.

  • Scroll to the temporary execution block at the bottom of retrieve.py.
  • Find this block:
if __name__ == "__main__":
    # Print the two counts that define the required benchmark input.
    ids, texts, qids, qtexts, labels = load()
    print(f"{len(ids)} documents")
    print(f"{len(qids)} queries")
  • Replace the temporary block with this dense-stage entry point:
def main():
    # Load the fixed benchmark and expose one dense ranking.
    ids, texts, qids, qtexts, labels = load()
    print(f"{len(ids)} documents")
    print(f"{len(qids)} queries")
    dense_rows = dense(texts, qtexts)
    print(f"dense first query top 5: {dense_rows[0][:5]}")


# Run the benchmark only when this file is executed directly.
if __name__ == "__main__":
    main()

How does the dense check work?

  • load() supplies the same corpus and 20 labelled queries used throughout the benchmark.
  • dense() returns 50 ranked corpus rows for each query.
  • The final print statement exposes the first five rows from the first dense ranking.
  • Save retrieve.py.

Expect the first dense run to take several minutes while your CPU encodes all 5,183 documents. The progress bar shows that the model is still working.

macOS

  • Run the dense stage from the active rerank-day environment by running:
python3 retrieve.py 2>&1 | tee -a check-out.txt

What does this command preserve?

The command runs the retrieval program. It appends the terminal evidence to check-out.txt.

Windows

  • Run the dense stage from the active rerank-day environment by running:
python retrieve.py 2>&1 | Tee-Object -Append check-out.txt

What does this command preserve?

The command runs the retrieval program. It appends the terminal evidence to check-out.txt.

Linux

  • Run the dense stage from the active rerank-day environment by running:
python3 retrieve.py 2>&1 | tee -a check-out.txt

What does this command preserve?

The command runs the retrieval program. It appends the terminal evidence to check-out.txt.

You will see 5183 documents followed by 20 queries. You will also see a line beginning with dense first query top 5: followed by five row indices.

Dense stage not finishing?

  • Confirm that the active terminal is inside the rerank-day folder.
  • Confirm that the virtual environment from Day Zero remains active.
  • Check that the cached dense model is available before retrying the command.

Still stuck? Help me diagnose my local dense retrieval run.

  • Record the working dense stage in the local repository by running these commands:
git add -A
git commit -m "Add dense search"

What does this commit capture?

The first command stages the dense implementation plus its local evidence. The second command records the stage as Add dense search.

The terminal confirms that the dense-stage commit was created. Your benchmark now has its first measurable search baseline.

Add sparse and hybrid retrieval

Sparse search ranks documents through lexical token overlap. Hybrid search combines the dense and sparse rankings through Reciprocal Rank Fusion.

  • In retrieve.py, add the sparse-search import and the existing rrf() import below the current import group:
import bm25s

from metrics import rrf

Why use these imports?

bm25s provides tokenization plus exact sparse retrieval. The imported rrf() function applies the fusion logic you already tested.

  • Add sparse() and hybrid() below dense() using this code:
def sparse(texts, qtexts):
    tokens = bm25s.tokenize(texts, stopwords="en")
    retriever = bm25s.BM25()
    retriever.index(tokens)
    results, scores = retriever.retrieve(bm25s.tokenize(qtexts, stopwords="en"), k=50)
    return results.tolist()


def hybrid(dense_rows, sparse_rows):
    return [rrf([dense_row, sparse_row])[:50] for dense_row, sparse_row in zip(dense_rows, sparse_rows)]

What do these search stages do?

  • The sparse stage tokenizes the corpus with English stopword handling before building its exact index.
  • Each sparse query returns 50 corpus row indices.
  • The hybrid stage passes each dense and sparse pair to rrf() before keeping the first 50 fused rows.
  • Rank fusion uses positions from each list. This avoids adding dense and sparse scores that use different scales.
  • Inside main(), add these lines immediately after the dense output:
    sparse_rows = sparse(texts, qtexts)
    print(f"sparse first query top 5: {sparse_rows[0][:5]}")
    hybrid_rows = hybrid(dense_rows, sparse_rows)
    print(f"hybrid first query top 5: {hybrid_rows[0][:5]}")

What will this reveal?

These lines run the lexical stage before fusing its ranking with the dense result. The two printed top-five lists make each new stage visible.

  • Save retrieve.py.

macOS

  • Run the expanded retrieval program and append its evidence by running:
python3 retrieve.py 2>&1 | tee -a check-out.txt

What does this run compare?

The program runs dense retrieval plus sparse retrieval over the same queries. It also prints the fused hybrid ranking for the first query.

Windows

  • Run the expanded retrieval program and append its evidence by running:
python retrieve.py 2>&1 | Tee-Object -Append check-out.txt

What does this run compare?

The program runs dense retrieval plus sparse retrieval over the same queries. It also prints the fused hybrid ranking for the first query.

Linux

  • Run the expanded retrieval program and append its evidence by running:
python3 retrieve.py 2>&1 | tee -a check-out.txt

What does this run compare?

The program runs dense retrieval plus sparse retrieval over the same queries. It also prints the fused hybrid ranking for the first query.

You will see lines beginning with sparse first query top 5: and hybrid first query top 5:. Each line contains five corpus row indices.

Missing sparse or hybrid output?

  • Check that import bm25s appears with the other third-party imports.
  • Check that from metrics import rrf appears before the constants.
  • Confirm that the new lines are indented inside main().

Need another pair of eyes? Help me debug the sparse and hybrid stages.

  • Record the sparse and hybrid stages in the local repository by running these commands:
git add -A
git commit -m "Add sparse and hybrid search"

What does this commit capture?

This snapshot preserves lexical retrieval plus rank fusion. Its commit message is Add sparse and hybrid search.

The repository now preserves three independently visible search stages. Dense supplies semantic evidence. Sparse supplies lexical evidence.

Rerank the candidates and test the metric guard

Cross-encoder reranking reads one query with each hybrid candidate before assigning a new relevance score. It can reorder the supplied 50 rows without adding a different document.

  • In retrieve.py, update the import area before adding the reranker model constant and rerank() with this code:
import time
from sentence_transformers import CrossEncoder, SentenceTransformer

RERANK_MODEL = "BAAI/bge-reranker-v2-m3"


def rerank(model, query, rows, texts):
    started = time.perf_counter()
    ranked = model.rank(query, [texts[row] for row in rows])
    milliseconds = (time.perf_counter() - started) * 1000
    reordered = [(rows[item["corpus_id"]], item["score"]) for item in ranked]
    return reordered, milliseconds

How does reranking stay measurable?

  • CrossEncoder loads BAAI/bge-reranker-v2-m3 for CPU scoring.
  • The timer measures the model's work in milliseconds for one query.
  • Each returned corpus_id points back to a position within the supplied candidate list.
  • The function returns reordered corpus rows with their scores. It also returns the measured time.
  • Inside main(), add these lines immediately after the hybrid output:
    model = CrossEncoder(RERANK_MODEL, max_length=512, device="cpu")
    ranked, milliseconds = rerank(model, qtexts[0], hybrid_rows[0], texts)
    print(f"hybrid+rerank first query top 5: {ranked[:5]}")
    print(f"first query rerank milliseconds: {milliseconds:.4f}")

What does the reranker check show?

The program creates one CPU reranker before scoring the first query's hybrid candidates. It prints the first five reordered rows with their scores.

The final line exposes the elapsed CPU time for that query.

  • Save retrieve.py.

macOS

  • Run all four search stages and append their evidence by running:
python3 retrieve.py 2>&1 | tee -a check-out.txt

What does this run establish?

The run reaches the reranker after dense and sparse retrieval produce the hybrid candidate set. Its final two lines expose the new ordering plus local CPU time.

Windows

  • Run all four search stages and append their evidence by running:
python retrieve.py 2>&1 | Tee-Object -Append check-out.txt

What does this run establish?

The run reaches the reranker after dense and sparse retrieval produce the hybrid candidate set. Its final two lines expose the new ordering plus local CPU time.

Linux

  • Run all four search stages and append their evidence by running:
python3 retrieve.py 2>&1 | tee -a check-out.txt

What does this run establish?

The run reaches the reranker after dense and sparse retrieval produce the hybrid candidate set. Its final two lines expose the new ordering plus local CPU time.

You will see a line beginning with hybrid+rerank first query top 5:. The next line begins with first query rerank milliseconds:.

Reranker output missing?

  • Confirm that the CrossEncoder import shares the existing Sentence Transformers import line.
  • Confirm that RERANK_MODEL matches the model identifier prepared during Day Zero.
  • Check that the reranker receives hybrid_rows[0] as its candidate rows.

If the run still stops early, help me debug the local reranker stage.

  • Record the working reranker in the local repository by running these commands:
git add -A
git commit -m "Add the reranker"

What does this commit preserve?

This commit captures the fourth search stage before the intentional scorer mutation. Its commit message is Add the reranker.

That completes the retrieval pipeline. Dense, sparse, hybrid, and hybrid plus rerank now produce local evidence from the same query set.

✔️ Awesome, I've got everything!

Great work. The completed retrieval stages are saved in retrieve.py and committed through Add the reranker.

ⓧ I'd like to double check the full code

Compare your complete retrieve.py with this reference:

import csv
import json
from pathlib import Path
import time

import bm25s
import numpy as np
from sentence_transformers import CrossEncoder, SentenceTransformer

from metrics import rrf

DATA = Path("data/scifact")
DENSE_MODEL = "BAAI/bge-small-en-v1.5"
RERANK_MODEL = "BAAI/bge-reranker-v2-m3"


def load():
    ids = []
    texts = []
    with (DATA / "corpus.jsonl").open(encoding="utf-8") as handle:
        for line in handle:
            row = json.loads(line)
            ids.append(str(row["_id"]))
            texts.append(f"{row.get('title', '')} {row.get('text', '')}".strip())

    query_text = {}
    with (DATA / "queries.jsonl").open(encoding="utf-8") as handle:
        for line in handle:
            row = json.loads(line)
            query_text[str(row["_id"])] = row["text"]

    all_labels = {}
    with (DATA / "qrels" / "test.tsv").open(encoding="utf-8") as handle:
        for row in csv.DictReader(handle, delimiter="\t"):
            qid = str(row["query-id"])
            corpus_id = str(row["corpus-id"])
            all_labels.setdefault(qid, set())
            if int(row["score"]) > 0:
                all_labels[qid].add(corpus_id)

    qids = list(all_labels)[:20]
    qtexts = [query_text[qid] for qid in qids]
    labels = {qid: all_labels[qid] for qid in qids}
    return ids, texts, qids, qtexts, labels


def dense(texts, qtexts):
    model = SentenceTransformer(DENSE_MODEL, device="cpu")
    cache = Path("corpus_emb.npy")
    if cache.exists():
        doc_emb = np.load(cache)
    else:
        doc_emb = model.encode(texts, normalize_embeddings=True, show_progress_bar=True)
        np.save(cache, doc_emb)
    query_emb = model.encode(qtexts, normalize_embeddings=True, show_progress_bar=False)
    return [np.argsort(-(doc_emb @ query))[:50].tolist() for query in query_emb]


def sparse(texts, qtexts):
    tokens = bm25s.tokenize(texts, stopwords="en")
    retriever = bm25s.BM25()
    retriever.index(tokens)
    results, scores = retriever.retrieve(bm25s.tokenize(qtexts, stopwords="en"), k=50)
    return results.tolist()


def hybrid(dense_rows, sparse_rows):
    return [rrf([dense_row, sparse_row])[:50] for dense_row, sparse_row in zip(dense_rows, sparse_rows)]


def rerank(model, query, rows, texts):
    started = time.perf_counter()
    ranked = model.rank(query, [texts[row] for row in rows])
    milliseconds = (time.perf_counter() - started) * 1000
    reordered = [(rows[item["corpus_id"]], item["score"]) for item in ranked]
    return reordered, milliseconds


def main():
    ids, texts, qids, qtexts, labels = load()
    print(f"{len(ids)} documents")
    print(f"{len(qids)} queries")
    dense_rows = dense(texts, qtexts)
    print(f"dense first query top 5: {dense_rows[0][:5]}")
    sparse_rows = sparse(texts, qtexts)
    print(f"sparse first query top 5: {sparse_rows[0][:5]}")
    hybrid_rows = hybrid(dense_rows, sparse_rows)
    print(f"hybrid first query top 5: {hybrid_rows[0][:5]}")
    model = CrossEncoder(RERANK_MODEL, max_length=512, device="cpu")
    ranked, milliseconds = rerank(model, qtexts[0], hybrid_rows[0], texts)
    print(f"hybrid+rerank first query top 5: {ranked[:5]}")
    print(f"first query rerank milliseconds: {milliseconds:.4f}")


if __name__ == "__main__":
    main()

A benchmark needs evidence that its scorer can reject a faulty metric definition. You will create that evidence without committing the defect.

  • In metrics.py, find ideal_hits = min(k, len(relevant)) inside ndcg_at_k().
  • Replace that line with ideal_hits = sum(1 for doc_id in ranked[:k] if doc_id in relevant).
  • Save metrics.py.

Before you run the scorer, do you think the nDCG fixture will still pass with a denominator based on retrieved relevant documents?

macOS

  • Run the intentionally modified scorer and append its evidence by running:
python3 metrics.py 2>&1 | tee -a check-out.txt

The Metric Guard Catches the Defect

You will see FAIL nDCG@5: got 0.6509, expected 0.4982 in the terminal. The command also appends that evidence to check-out.txt.

Linux users can run the same command.

Windows

  • Run the intentionally modified scorer and append its evidence by running:
python metrics.py 2>&1 | Tee-Object -Append check-out.txt

The Metric Guard Catches the Defect

You will see FAIL nDCG@5: got 0.6509, expected 0.4982 in the terminal. PowerShell also appends that evidence to check-out.txt.

The failure is the intended result. It proves that the fixture detects this incorrect nDCG denominator before a misleading score reaches the evaluation.

  • Leave the modified metrics.py line uncommitted for the next step.

You now have four committed search stages plus preserved proof that the metric guard catches the intended defect. Next up, you will restore the scorer before measuring every row and closing out the evidence review.

Evaluate and Tag the Review

Your retrieval pipeline now produces dense, sparse, hybrid, and reranked evidence. The preserved metric failure also proves that your scorer catches a broken ideal denominator.

This final step restores the scorer before measurement. You will evaluate every method against the same human relevance labels. You will then trace each shipped claim before preserving the review with a local Git tag.

In this step, get ready to:
  • Restore the scorer to its passing state.
  • Evaluate all four retrieval methods against the same labels.
  • Trace the close-out artifacts before tagging the release.
Restore the metric guard

The intended failure remains preserved in check-out.txt. The temporary denominator in metrics.py now has to be corrected before any measured quality claim is trusted.

  • In metrics.py, find this temporary line:
ideal_hits = sum(1 for doc_id in ranked[:k] if doc_id in relevant)
  • Replace the temporary calculation with the corrected line shown below:
ideal_hits = min(k, len(relevant))

Why does the restored denominator work?

The temporary calculation counted only relevant documents that already appeared in the measured ranking. A weaker ranking could therefore shrink its own ideal score, which produced 0.6509 instead of 0.4982 for the fixture.

The restored calculation uses the smaller of the cutoff or the total number of relevant documents. The ideal top-five score now stays independent of the ranking being evaluated.

  • Save metrics.py.

Before you rerun the scorer, do you expect the same fixture to pass or fail with the restored denominator?

macOS

  • Append the restored scorer result to check-out.txt by running:
python3 metrics.py 2>&1 | tee -a check-out.txt

Windows

  • Append the restored scorer result to check-out.txt by running:
python metrics.py 2>&1 | Tee-Object -Append check-out.txt

Linux

  • Append the restored scorer result to check-out.txt by running:
python3 metrics.py 2>&1 | tee -a check-out.txt

You'll see three PASS lines followed by ALL CHECKS PASS. The same log now preserves the caught defect and the successful restoration.

The metric guard is working again. One source check remains: confirm the restored file matches the committed scorer.

Before you inspect the source difference, do you expect Git to find a remaining change in metrics.py?

  • Check that metrics.py matches its committed version by running:
git diff -- metrics.py

You'll see no output. That empty result proves metrics.py matches its committed passing version.

Still seeing a scorer difference?

  • Check that the restored line uses min(k, len(relevant)) exactly.
  • Check the indentation against the surrounding ndcg_at_k() function.
  • Save metrics.py before running the difference check again.

Need another pair of eyes? Help me restore the nDCG denominator without losing the preserved failure evidence.

Build and run the evaluator

The evaluator turns the retrieval stages into one comparable decision record. It scores the same first 20 queries with Recall@5, Recall@50, and nDCG@5.

It also writes the measured rows to JSON. The output includes signed changes from dense, same-candidate-set evidence, and measured p95 rerank time.

  • Create evaluate.py inside the rerank-day folder using the editor from earlier.
  • Add the imports and scoring helpers by copying this first block into evaluate.py:
import json

import numpy as np
from sentence_transformers import CrossEncoder

from metrics import ndcg_at_k, recall_at_k
from retrieve import RERANK_MODEL, dense, hybrid, load, rerank, sparse


def mean(values):
    return sum(values) / len(values)


def score_run(run_rows, ids, qids, labels):
    recall5 = []
    recall50 = []
    ndcg5 = []
    for rows, qid in zip(run_rows, qids):
        ranked_ids = [ids[row] for row in rows]
        recall5.append(recall_at_k(ranked_ids, labels[qid], 5))
        recall50.append(recall_at_k(ranked_ids, labels[qid], 50))
        ndcg5.append(ndcg_at_k(ranked_ids, labels[qid], 5))
    return {
        "recall@5": mean(recall5),
        "recall@50": mean(recall50),
        "nDCG@5": mean(ndcg5),
    }

What do the scoring helpers do?

  • The imports reuse the metric functions and retrieval stages you already tested.
  • The mean() helper averages one metric across all 20 queries.
  • The score_run() helper converts row positions into SciFact document IDs before scoring them against the human labels.
  • Save evaluate.py.
  • Confirm the scoring helpers are present in the editor.

Seeing an incomplete helper?

  • Check that score_run() returns all three metric fields.
  • Check that each metric list is populated inside the query loop.

If the structure still looks different, help me compare my scoring helpers with the supplied evaluator structure.

  • Append the evaluation setup below the scoring helpers in evaluate.py:
def main():
    ids, texts, qids, qtexts, labels = load()
    dense_rows = dense(texts, qtexts)
    sparse_rows = sparse(texts, qtexts)
    hybrid_rows = hybrid(dense_rows, sparse_rows)
    model = CrossEncoder(RERANK_MODEL, max_length=512, device="cpu")

    reranked_rows = []
    reranked_scores = []
    rerank_ms = []
    for query, rows in zip(qtexts, hybrid_rows):
        ranked, milliseconds = rerank(model, query, rows, texts)
        reranked_rows.append([row for row, score in ranked])
        reranked_scores.append([score for row, score in ranked])
        rerank_ms.append(milliseconds)

How is the comparison kept fair?

The same corpus, query texts, query IDs, and labels feed all four methods. The CrossEncoder receives only each hybrid method's 50 candidates.

The reranking loop stores candidate order, model scores, and elapsed milliseconds separately. That makes quality and CPU time visible in the same evaluation.

  • Save evaluate.py.
  • Confirm the reranking loop contains the row order, scores, and timing lists.

Missing part of the reranking loop?

  • Keep the loop indented inside main().
  • Check that rerank() receives the model, query, hybrid rows, and corpus texts.

If the loop remains unclear, help me compare the reranking loop with the supplied evaluator.

  • Append the four-row scoring and comparison logic inside main():
    runs = {
        "dense": dense_rows,
        "sparse": sparse_rows,
        "hybrid": hybrid_rows,
        "hybrid+rerank": reranked_rows,
    }
    scored = {name: score_run(rows, ids, qids, labels) for name, rows in runs.items()}
    dense_score = scored["dense"]
    table = []
    for name in ("dense", "sparse", "hybrid", "hybrid+rerank"):
        row = {
            "search": name,
            **scored[name],
            "delta_recall@5_vs_dense": scored[name]["recall@5"] - dense_score["recall@5"],
            "delta_nDCG@5_vs_dense": scored[name]["nDCG@5"] - dense_score["nDCG@5"],
        }
        table.append(row)
        print(json.dumps(row))

    same_50 = all(set(before) == set(after) for before, after in zip(hybrid_rows, reranked_rows))
    p95_ms = float(np.percentile(rerank_ms, 95))
    print(f"same 50 rows after rerank: {'yes' if same_50 else 'no'}")
    print(f"p95 rerank milliseconds: {p95_ms:.4f}")

What does the report measure?

  • The runs mapping gives each retrieval method one scored row.
  • The two delta fields subtract the dense baseline while retaining positive or negative direction.
  • The same_50 check compares candidate membership before and after reranking.
  • The p95_ms value captures the 95th percentile of measured per-query reranking time.
  • Save evaluate.py.
  • Confirm the four method names match the retrieval comparison exactly.

Do the report fields look different?

  • Check that the dense row remains the baseline for both signed changes.
  • Check that same_50 compares sets from the hybrid and reranked rows.

If a field is missing, help me compare the report fields with the supplied evaluator.

  • Append the result-file writer and main guard at the bottom of evaluate.py:
    results = {
        "documents": len(ids),
        "queries": len(qids),
        "table": table,
        "same_50_rows_after_rerank": same_50,
        "p95_rerank_milliseconds": p95_ms,
        "reranked_top_5": {
            qid: [
                {"corpus_id": ids[row], "score": score}
                for row, score in zip(rows[:5], scores[:5])
            ]
            for qid, rows, scores in zip(qids, reranked_rows, reranked_scores)
        },
    }
    with open("results.json", "w", encoding="utf-8") as handle:
        json.dump(results, handle, indent=2)


if __name__ == "__main__":
    main()

What goes into results.json?

The result file records the corpus and query counts. It also stores the four metric rows, same-50 outcome, p95 timing, and each query's reranked top five.

The main guard runs the evaluation when you execute evaluate.py directly.

  • Save evaluate.py.

Before you run the evaluator, do you expect reranking to preserve all 50 hybrid candidates?

This CPU evaluation is the slowest check in the step. A quiet terminal while the reranker processes the 20 queries is expected.

  • Run the evaluation and append its output to check-out.txt using the command for your platform:

macOS

python3 evaluate.py 2>&1 | tee -a check-out.txt

Windows

python evaluate.py 2>&1 | Tee-Object -Append check-out.txt

Linux

python3 evaluate.py 2>&1 | tee -a check-out.txt

You should see four JSON rows named dense, sparse, hybrid, and hybrid+rerank.

You should then see the same-50 result and the measured p95 rerank milliseconds. The run also creates results.json with the exact local measurements.

Evaluator stopped before writing results?

  • Confirm the active terminal still uses the .venv from earlier.
  • Compare every imported function name with metrics.py and retrieve.py.
  • Check that each appended block remains indented inside main() where shown.

If the run still stops, help me diagnose my evaluator without changing the benchmark design.

  • Copy the measured dense Recall@5 value into Decision Record 1 in docs/decisions.md.
  • Copy the measured dense Recall@50 value into the same decision record.
  • Copy the measured dense nDCG@5 value into the same decision record.
  • Save docs/decisions.md.
  • Stage the evaluator evidence by running:
git add -A

This stages evaluate.py, results.json, the updated validation log, and the measured decision record.

  • Commit the evaluator task by running:
git commit -m "Add the evaluator"

You should see a new commit named Add the evaluator. Your measured comparison now has a fixed point in local history.

✔️ Awesome, I've got everything!

Strong work. Your four-row evaluation is committed with the exact local measurements in results.json.

ⓧ I'd like to double check the full code

Compare your saved evaluate.py with this complete reference.

import json

import numpy as np
from sentence_transformers import CrossEncoder

from metrics import ndcg_at_k, recall_at_k
from retrieve import RERANK_MODEL, dense, hybrid, load, rerank, sparse


def mean(values):
    return sum(values) / len(values)


def score_run(run_rows, ids, qids, labels):
    recall5 = []
    recall50 = []
    ndcg5 = []
    for rows, qid in zip(run_rows, qids):
        ranked_ids = [ids[row] for row in rows]
        recall5.append(recall_at_k(ranked_ids, labels[qid], 5))
        recall50.append(recall_at_k(ranked_ids, labels[qid], 50))
        ndcg5.append(ndcg_at_k(ranked_ids, labels[qid], 5))
    return {
        "recall@5": mean(recall5),
        "recall@50": mean(recall50),
        "nDCG@5": mean(ndcg5),
    }


def main():
    ids, texts, qids, qtexts, labels = load()
    dense_rows = dense(texts, qtexts)
    sparse_rows = sparse(texts, qtexts)
    hybrid_rows = hybrid(dense_rows, sparse_rows)
    model = CrossEncoder(RERANK_MODEL, max_length=512, device="cpu")

    reranked_rows = []
    reranked_scores = []
    rerank_ms = []
    for query, rows in zip(qtexts, hybrid_rows):
        ranked, milliseconds = rerank(model, query, rows, texts)
        reranked_rows.append([row for row, score in ranked])
        reranked_scores.append([score for row, score in ranked])
        rerank_ms.append(milliseconds)

    runs = {
        "dense": dense_rows,
        "sparse": sparse_rows,
        "hybrid": hybrid_rows,
        "hybrid+rerank": reranked_rows,
    }
    scored = {name: score_run(rows, ids, qids, labels) for name, rows in runs.items()}
    dense_score = scored["dense"]
    table = []
    for name in ("dense", "sparse", "hybrid", "hybrid+rerank"):
        row = {
            "search": name,
            **scored[name],
            "delta_recall@5_vs_dense": scored[name]["recall@5"] - dense_score["recall@5"],
            "delta_nDCG@5_vs_dense": scored[name]["nDCG@5"] - dense_score["nDCG@5"],
        }
        table.append(row)
        print(json.dumps(row))

    same_50 = all(set(before) == set(after) for before, after in zip(hybrid_rows, reranked_rows))
    p95_ms = float(np.percentile(rerank_ms, 95))
    print(f"same 50 rows after rerank: {'yes' if same_50 else 'no'}")
    print(f"p95 rerank milliseconds: {p95_ms:.4f}")

    results = {
        "documents": len(ids),
        "queries": len(qids),
        "table": table,
        "same_50_rows_after_rerank": same_50,
        "p95_rerank_milliseconds": p95_ms,
        "reranked_top_5": {
            qid: [
                {"corpus_id": ids[row], "score": score}
                for row, score in zip(rows[:5], scores[:5])
            ]
            for qid, rows, scores in zip(qids, reranked_rows, reranked_scores)
        },
    }
    with open("results.json", "w", encoding="utf-8") as handle:
        json.dump(results, handle, indent=2)


if __name__ == "__main__":
    main()
Trace the close-out and tag the release

Measured output supports a claim only when the wording stays within what the design evidence or local run establishes. The close-out uses your existing AI client for drafting while you remain responsible for every shipped statement.

The trace review covers factual, causal, evaluative, performance, and recommendation claims. Unsupported claims must be removed or visibly struck before the release is tagged.

  • Copy the exact four JSON rows from the evaluator output.
  • Copy the exact same-50 line from the evaluator output.
  • Copy the exact p95 line from the evaluator output.
  • Switch back to the AI client from earlier.
  • Attach or paste the contents of docs/requirements-brief.md, docs/decisions.md, docs/topology.md, docs/build-plan.md, and docs/source-map.md so the AI client can use the five design documents as sources.
  • Confirm the AI client can read all five documents before requesting the drafts.
  • Paste the following request into the AI client:
Title the teach-back and readout exactly "The reranker that has to pay for itself". Draft docs/TEACHBACK.md, docs/AAR.md, and one self-contained readout.html. Use only the five design documents and the pasted run result as claim sources.

The teach-back must explain what dense, sparse, hybrid, and hybrid plus rerank each do; why reranking cannot change the 50 candidates; when reranking is not worth its time; and why 20 queries prove no general win.

The after-action review must answer expected, happened, why, and next; say whether the measured lift surprised the learner; state that 20 queries prove no general win; and reserve a line for the first-draft trace score as untraceable claims over total claims.

The readout must name the four rows, show every signed change over dense and p95 rerank time, and include this sourced context: WHO projected a global shortage of 11.1 million health workers by 2030, source https://www.who.int/news/item/18-09-2026-one-in-four-doctors-nearing-retirement-age-as-who-warns-of-global-health-workforce-gap, dated 18 September 2026. It must say this context does not make SciFact clinical-guideline evidence.

The readout provenance must also state the dated implementation reason: Sentence Transformers v6.0.0, published 18 August 2026, made CrossEncoder.rank score values Python floats that are directly JSON serializable. Cite https://github.com/huggingface/sentence-transformers/releases/tag/v6.0.0.

Do not invent, estimate, round differently, or generalize any measured value. Do not modify metrics.py, retrieve.py, evaluate.py, results.json, or check-out.txt.

Why is the drafting prompt strict?

The request limits claim sources to the frozen design documents and pasted run evidence. It also protects the implementation files and measured outputs from AI edits.

The required limitations keep the 20-query SciFact result within its real scope. The readout can support a local retrieval decision without implying a general or clinical win.

  • Paste the copied run output beneath the request.
  • Submit the complete request to generate the three first drafts.
  • Use your editor's new-file control to create docs/TEACHBACK.md.
  • Paste the teach-back draft into docs/TEACHBACK.md.
  • Use your editor's new-file control to create docs/AAR.md.
  • Paste the after-action review draft into docs/AAR.md.
  • Use your editor's new-file control to create readout.html inside rerank-day.
  • Paste the HTML draft into readout.html.

Where are you now?

The three drafts now exist locally. They remain draft evidence until every claim has a traceable source.

  • Save docs/TEACHBACK.md.
  • Save docs/AAR.md.
  • Save readout.html.
  • Count every factual claim across all three drafts.
  • Count every causal claim across all three drafts.
  • Count every evaluative claim across all three drafts.
  • Count every performance claim across all three drafts.
  • Count every recommendation claim across all three drafts.
  • Classify each counted claim as traceable only when a design document or pasted run result supports it.
  • Replace the reserved trace-score line in docs/AAR.md with your actual counts using this format:
First-draft trace score: untraceable claims / total claims = U/T

How should the trace score be completed?

Replace U with the number of unsupported first-draft claims. Replace T with the total number of claims you counted.

Keep the first-draft fraction after cleaning the artifacts. It records how much unsupported material the review removed.

  • Save docs/AAR.md.
  • Confirm the completed trace-score line shows your actual fraction.

Unsure whether a claim is traceable?

  • Match measured claims against the exact evaluator output.
  • Match design claims against one of the five supplied documents.
  • Classify unsupported generalizations as untraceable.

If a claim still feels ambiguous, help me test whether this claim is supported by my supplied evidence.

  • Remove or strike every untraceable claim from docs/TEACHBACK.md.
  • Remove or strike every untraceable claim from docs/AAR.md.
  • Remove or strike every untraceable claim from readout.html.
  • Confirm docs/TEACHBACK.md explains all four methods.
  • Confirm the teach-back explains same-50 reranking.
  • Confirm the teach-back explains the CPU tradeoff.
  • Confirm the teach-back retains the 20-query limit.
  • Confirm docs/AAR.md answers expected, happened, why, and next.

What has the trace review proved?

Every remaining statement now has a path back to the frozen design or measured run. The artifacts can explain the decision without borrowing confidence from unsupported wording.

  • Confirm docs/AAR.md states whether the measured lift surprised you.
  • Confirm the exact title appears in docs/TEACHBACK.md.
  • Open readout.html from the rerank-day folder in your browser.

You should see a self-contained page titled The reranker that has to pay for itself. It should show the exact run rows, signed changes, p95 timing, sourced context, provenance, and evidence limits without a server.

Does the readout depend on something remote?

  • Remove any stylesheet reference that points outside readout.html.
  • Move required styling into the local HTML file.
  • Replace any missing remote image with self-contained text or markup.

Need help making the page portable? Help me identify remote dependencies in my local readout without changing its claims.

  • Stage the traced close-out artifacts by running:
git add -A

This stages the reviewed teach-back, after-action review, self-contained readout, and their trace corrections.

  • Commit the reviewed close-out by running:
git commit -m "Add the readout"

You should see a new commit named Add the readout.

  • Create the local close-out tag by running:
git tag v1.0-rerank

The tag now points to the reviewed 20-query release. It preserves the required close-out before any optional population expansion.

Before you list the tags, which release name do you expect Git to return?

  • Verify the local release tag by running:
git tag --list

You should see v1.0-rerank in the output. That is the review closed out with measured results, preserved validation evidence, and traced claims.

Tag missing from the list?

  • Confirm the readout commit completed before you created the tag.
  • Create the tag again from the current close-out commit.
  • Run the tag-list check once more.

If the tag still does not appear, help me verify my local Git tag without adding a remote.

Secret mission

Score All 300 Test Queries

Your tagged release measures 20 human-labelled queries. Extend the same four-method evaluation to all 300 SciFact test queries, then check whether both signed quality changes stay positive.

Clean Up Your Resources

Clean Up Your Resources

Your benchmark stays on your machine at an ongoing cost of $0.00. Decide whether to keep the evidence, pause the session, or delete the local folder.

Resources you used:

  • The local rerank-day Git repository containing the three Python programs plus five design documents.
  • The local runtime files inside rerank-day, including .venv/, data/, plus corpus_emb.npy.
  • The evaluation evidence inside rerank-day, including check-out.txt, results.json, plus results-300.json.
  • The traced close-out artifacts inside rerank-day, including docs/TEACHBACK.md, docs/AAR.md, plus readout.html.
  • The local v1.0-rerank tag preserving the original 20-query release.

Keep everything running

No action is needed. Choose this if you want to inspect the tagged release or repeat either evaluation.

  • Keep rerank-day to preserve the 20-query release plus the optional 300-query result.
  • Retain the package cache plus model cache to make future runs faster.
  • Use v1.0-rerank whenever you need the unchanged pre-mission release.

The evaluation has finished, so the terminal does not need to remain open.

Pause - I'll come back to this later

Closing your local tools ends the session while preserving every result. No service remains active in the background.

  • Save any open project files in your editor.
  • Close your editor.
  • Close the terminal used for the 300-query evaluation.
  • Leave rerank-day in its current location.
  • Leave the separate package cache plus model cache in place for future runs.

Delete - I don't want to use this again

Deleting the project folder is permanent. The tagged release plus both measured result files disappear with it.

The deletion path depends on your operating system.

macOS

  • Copy any evidence you want to keep into another folder outside rerank-day.
  • Close every Terminal session currently using rerank-day.
  • Close every editor window currently using files from rerank-day.
  • Use Finder to locate the parent folder containing rerank-day.
  • Move rerank-day to the Trash.
  • Confirm that rerank-day no longer appears in the parent folder.
  • Empty the Trash if you want to reclaim the disk space immediately.

Windows

  • Copy any evidence you want to keep into another folder outside rerank-day.
  • Close every PowerShell session currently using rerank-day.
  • Close every editor window currently using files from rerank-day.
  • Use File Explorer to locate the parent folder containing rerank-day.
  • Move rerank-day to the Recycle Bin.
  • Confirm that rerank-day no longer appears in the parent folder.
  • Empty the Recycle Bin if you want to reclaim the disk space immediately.

Linux

  • Copy any evidence you want to keep into another folder outside rerank-day.
  • Close every terminal session currently using rerank-day.
  • Close every editor window currently using files from rerank-day.
  • Use your desktop file manager to locate the parent folder containing rerank-day.
  • Move rerank-day to your desktop trash.
  • Confirm that rerank-day no longer appears in the parent folder.
  • Empty your desktop trash if you want to reclaim the disk space immediately.

Keep Shared Caches Safe

The separate model cache can serve other local projects. This cleanup leaves it in place because the project records no safe shared-cache path.

  • Keep the model cache unless you independently confirm its location.
  • Confirm that no other local project uses the model cache before deleting it.

Nice Work!

Nice Work!

You did it! You built a tagged local retrieval decision review that tests whether cross-encoder reranking earns its measured CPU cost. The review ties every quality claim to human relevance labels or recorded run evidence.

You've learned how to:

  • Built a local benchmark over 5,183 SciFact documents. Scored dense search, sparse search, hybrid search, and hybrid plus rerank on the same 20 queries.
  • Proved the scorer catches an intentionally broken nDCG@5 denominator. Restored the denominator before evaluation. Measured Recall@5, Recall@50, signed change over dense, and p95 rerank time.
  • Shipped a claim-traced teach-back plus an after-action review. Produced a self-contained HTML readout from the measured evidence. Preserved the 20-query release under the local v1.0-rerank tag.
  • Completed the Secret Mission by scoring all 300 SciFact test queries in results-300.json. Used both signed quality changes to record whether the reranker lift held beyond the original sample.

Ready to quiz yourself?