Build a Mini-Shazam
Build a noise-tolerant audio recognizer with hash lookup and offset voting.
Introduction
30 Second Summary
Background noise rarely stops you from recognizing a familiar song. Starting halfway through the track usually does not fool you either.
In this project, you will build a local music recognizer using landmark audio fingerprints. It identifies a shifted noisy excerpt from a three-track catalog while showing the evidence behind each match.
What You'll Build
Your finished recognizer identifies Vibe Ace from a shifted noisy clip while showing why the match won.
By the end of this project, you'll have:
- Clear failure proof showing a direct waveform matcher reject the correct recording once the excerpt starts later in the track.
- A constellation map you can open in constellation.png to see cyan peaks marking the strongest parts of a spectrogram.
- A tested fingerprint search that identifies known clips across controlled offsets plus noise levels. A reproducible benchmark reports accuracy, false positives, feature cost, plus lookup latency. A production diagram shows how an inverted index scales across fingerprint workers plus index shards.
- Secret Mission: Add the out-of-catalog choice recording. Calibrate aligned-vote plus confidence thresholds until the recognizer returns No confident match.
Are there any prerequisites?
You need a Mac with Visual Studio Code. You should be familiar with basic Python plus system design. No audio-processing experience is required.
Before We Start
Before any hands-on work, lock in the system you are about to build. Your explainable recognizer will identify a shifted, noisy catalog excerpt through an inverted index that retrieves candidate hash postings without scanning every full recording.
Set Up the Reproducible Audio Lab
Every benchmark result depends on the interpreter plus the package versions that produced it. Your audio lab needs a controlled Python environment before you compare matching methods.
A virtual environment isolates this project's libraries. Pinned versions make every later result easier to reproduce.
In this step, get ready to:
- Confirm that your Mac has Python 3.12 or newer.
- Create the Visual Studio Code workspace with the complete file structure.
- Install the pinned packages inside an isolated .venv environment.
Confirm Python 3.12 or newer
The pinned scientific packages require Python 3.12 or newer. This check determines whether your existing interpreter can create the lab.
- Press Cmd+Space to open Spotlight Search.
- Open Terminal by typing Terminal into Spotlight Search before pressing Enter.
- Check your installed Python version by running this command:
python3 --version
What does this version check prove?
The python3 executable reports the interpreter available in your terminal. A result of Python 3.12 or newer supports every package pin in this project.
Choose the tab that matches the version result in your terminal.
✔️ I see version 3.12 or higher
Your installed interpreter meets the requirement. Keep Terminal open for the project setup.
ⓧ I see an older version
The installed interpreter cannot use every pinned package in this project. A current macOS installer gives you a compatible Python release.
- Open the official Python downloads page for macOS in your browser.
- Download a current macOS installer that provides Python 3.12 or newer.
- Run the downloaded installer.
- Complete the installation prompts.
- Quit Terminal after the installation finishes.
- Press Cmd+Space to reopen Spotlight Search.
- Reopen Terminal by typing Terminal before pressing Enter.
- Recheck the installed version by running:
python3 --version
What should the new version show?
The new terminal session reloads your command paths. Continue when the result reports Python 3.12 or newer.
ⓧ Command not found
Your terminal cannot currently locate Python 3. Installing a current release adds the interpreter required by this project.
- Open the official Python downloads page for macOS in your browser.
- Download a current macOS installer that provides Python 3.12 or newer.
- Run the downloaded installer.
- Complete the installation prompts.
- Quit Terminal after the installation finishes.
- Press Cmd+Space to reopen Spotlight Search.
- Reopen Terminal by typing Terminal before pressing Enter.
- Confirm that Terminal can now locate Python by running:
python3 --version
What confirms the installation?
A printed Python version confirms that Terminal can locate the interpreter. Continue when that version is 3.12 or newer.
Create the Visual Studio Code project files
A folder opened in Visual Studio Code becomes the workspace for your files plus its interpreter choice. You will reserve every source path now without adding recognizer code from later steps.
- Press Cmd+Space to open Spotlight Search.
- Open Visual Studio Code by typing Visual Studio Code before pressing Enter.
- Select File from the top menu.
- Select Open Folder.
- Navigate to your Desktop in the folder picker.
- Select New Folder.
- Enter mini-shazam as the folder name.
- Select Open to load the folder as your workspace.
You should see mini-shazam at the top of the Explorer view. That folder now holds the complete local project.
The dependency manifest records exact package versions. These pins keep the audio pipeline plus its plots consistent across later runs.
- Select the Explorer view in the Activity Bar.
- Select the New File control in the Explorer toolbar.
- Enter requirements.txt before pressing Enter.
- Add the pinned dependency manifest by pasting this exact file content:
librosa==1.0.0
matplotlib==3.11.2
numpy==2.5.4
scipy==1.18.1
What do these packages provide?
- The librosa package loads audio before turning it into time-frequency data.
- Matplotlib saves the constellation plot you will build later.
- NumPy handles the numerical arrays used for audio samples.
- SciPy provides the local maximum filter used for peak detection.
Each == pin selects one package release. This file acts like a controlled build artifact for your local experiment.
- Save requirements.txt by pressing Cmd+S.
- Confirm that requirements.txt remains visible under mini-shazam in the Explorer view.
Does the dependency file look different?
- Make sure each package pin appears on its own line.
- Remove any extra spaces around each == symbol.
Ask for help comparing the package pins in requirements.txt.
✔️ Awesome, I've got everything!
Your dependency manifest is saved with all four package pins. Continue with the placeholder source files.
ⓧ I'd like to double check the full code
Compare your saved requirements.txt with this complete file.
librosa==1.0.0
matplotlib==3.11.2
numpy==2.5.4
scipy==1.18.1
What should match?
The four package pins plus their order should match exactly. Save the corrected file before continuing.
The remaining files reserve the final project structure. Each file stays empty until its assigned build step.
- Create an empty recognizer.py file with the Explorer toolbar's New File control.
- Create an empty evaluate.py file with the same control.
You should now see recognizer.py plus evaluate.py in the Explorer view.
- Create an empty test_recognizer.py file with the Explorer toolbar's New File control.
- Create an empty README.md file with the same control.
You should now see the test path plus the project README under mini-shazam.
- Create an empty threshold_recognizer.py file with the Explorer toolbar's New File control.
- Create an empty evaluate_unknown.py file with the same control.
You should now see both threshold-related source paths in the Explorer view.
- Create an empty test_thresholds.py file with the Explorer toolbar's New File control.
Your workspace should contain requirements.txt plus all seven empty project files. No audio has been downloaded at this point.
Install and verify the isolated environment
The integrated terminal starts inside the open mini-shazam workspace. Commands run there can create the environment beside your project files.
- Select View from the top menu.
- Select Terminal to open the integrated terminal.
- Create the isolated environment by running this command:
python -m venv .venv
What does this command create?
The venv module creates a .venv folder inside your workspace. That folder stores a project-specific interpreter plus its installed packages.
- Confirm that .venv appears under mini-shazam in the Explorer view.
Can't create the virtual environment?
- Confirm that the integrated terminal belongs to the open mini-shazam workspace.
- Recheck that your installed Python version is 3.12 or newer.
Ask for help creating the .venv environment.
- Activate the new environment by running:
source .venv/bin/activate
What does activation change?
Activation places the environment's Python executable first for this terminal session. Package installations now stay inside .venv.
Your terminal prompt should now begin with (.venv). That prefix confirms the isolated environment is active.
Don't see the environment prefix?
- Confirm that the .venv folder exists in the Explorer view.
- Run the activation command again from the integrated terminal.
Ask for help activating .venv.
The first dependency install may take a few minutes while package files download. The terminal remains usable when its prompt returns.
- Install every pinned dependency into the active environment by running:
python -m pip install -r requirements.txt
How does the installation stay reproducible?
- The pip installer reads every pin from requirements.txt.
- The active environment directs those packages into .venv.
- The fixed versions control the software inputs used by every later benchmark.
The installation succeeds when the terminal prompt returns without an installation error. Your isolated environment now contains all four pinned packages.
Did the package installation fail?
- Confirm that your terminal prompt starts with (.venv).
- Compare every line in requirements.txt with the complete file above.
- Confirm that the Mac has an active internet connection for package downloads.
Ask for help installing the pinned dependencies.
Visual Studio Code must use the same interpreter as your terminal. Selecting .venv keeps future runs plus editor analysis attached to the pinned environment.
- Press Cmd+Shift+P to open the Command Palette.
- Enter Python: Select Interpreter into the Command Palette.
- Select the interpreter whose path ends with .venv/bin/python.
The Status Bar should now show an interpreter from .venv. This selection connects future Python runs to the environment you just prepared.
Can't find the .venv interpreter?
- Wait a few seconds for Visual Studio Code to finish environment discovery.
- Run Python: Select Interpreter again from the Command Palette.
- Install Microsoft's Python extension from the Extensions view if the selection command is missing.
Ask for help finding the environment interpreter.
Before you run the final check, which folder do you expect Matplotlib to report as its installation location?
- Verify Matplotlib's version plus its installation location by running:
python -c 'import matplotlib; print(matplotlib.__version__, matplotlib.__file__)'
What does this verification inspect?
- The matplotlib.__version__ value reports the installed release.
- The matplotlib.__file__ value reports the package location used by this interpreter.
You should see 3.11.2 followed by a path containing .venv. That's your reproducible lab ready: Matplotlib now comes from the isolated environment you control.
Seeing the wrong version or path?
- Confirm that the terminal prompt begins with (.venv).
- Reactivate the environment from the integrated terminal.
- Reinstall the pins from requirements.txt after activation.
Ask for help fixing the Matplotlib verification.
Your project structure plus its pinned environment are ready. Next, you will build a direct waveform matcher and watch a shifted noisy query expose its alignment problem.
Expose the Baseline Failure
Your Python environment is ready. The next question is whether raw audio samples provide a reliable way to identify a recording that starts at a different point.
A direct waveform template depends on sample alignment. In this step, you will use librosa to load a small catalog before testing that assumption with a shifted noisy query.
In this step, get ready to:
- Load three licensed recordings into a deterministic local catalog.
- Create a shifted two-second query with reproducible Gaussian noise.
- Run a strict opening-template comparison to expose its alignment failure.
Load the licensed catalog
The baseline needs a small reference catalog with consistent audio settings. A shared sample rate keeps every recording comparable.
- Select recognizer.py in the Explorer sidebar in Visual Studio Code.
- Replace the placeholder contents with the imports and catalog configuration below:
from __future__ import annotations
import math
from dataclasses import dataclass
import librosa
import numpy as np
SAMPLE_RATE = 11_025
CATALOG_SECONDS = 8.0
QUERY_SECONDS = 2.0
QUERY_SAMPLES = int(SAMPLE_RATE * QUERY_SECONDS)
BASELINE_ACCEPT_SCORE = 0.80
CATALOG_KEYS = ("brahms", "vibeace", "trumpet")
CATALOG_LABELS = {
"brahms": "Hungarian Dance number 5",
"vibeace": "Vibe Ace",
"trumpet": "solo trumpet 06",
}
What does this configuration control?
- The imports provide audio loading, array operations, mathematical helpers, and data-class support.
- The catalog uses eight seconds from each recording at 11_025 samples per second.
- The query contains two seconds of audio through QUERY_SAMPLES.
- The baseline accepts a result only when its similarity reaches 0.80.
- Save recognizer.py with Cmd+S.
The unsaved dot disappears from the recognizer.py tab. Your catalog configuration is now stored.
Seeing import warnings?
Confirm that the selected interpreter points inside .venv. The pinned packages were installed in that environment.
If the warning remains, help me check the interpreter used by recognizer.py.
Each catalog key asks librosa for one documented example recording. The first request downloads the audio before later runs use the local cache.
- Add the catalog loader and the first version of main() below CATALOG_LABELS by pasting this code:
def load_catalog() -> dict[str, object]:
catalog: dict[str, object] = {}
for key in CATALOG_KEYS:
path = librosa.example(key)
audio, _ = librosa.load(
path,
sr=SAMPLE_RATE,
mono=True,
duration=CATALOG_SECONDS,
)
catalog[key] = audio
return catalog
def main() -> None:
print("Loading the licensed reference catalog...")
catalog = load_catalog()
if __name__ == "__main__":
main()
How does the catalog load?
- librosa.example() resolves each documented example recording.
- librosa.load() converts each recording to mono at the shared sample rate.
- duration=CATALOG_SECONDS limits every reference to the first eight seconds.
- main() starts the catalog load when you run the file.
- Save recognizer.py.
The first run may pause for a short while as librosa downloads the three recordings. Later runs use the cached copies.
- Load the catalog from the active VS Code terminal by running:
python recognizer.py
What does this run prove?
The script resolves all three catalog keys before returning control to the terminal. This confirms that the reference audio can be loaded with the pinned environment.
You will see Loading the licensed reference catalog... before the terminal returns to its prompt. Good, your three-track reference catalog is available for the experiment.
Which recordings are in the catalog?
- brahms is “Hungarian Dance number 5” by Johannes Brahms, performed by the US Army Strings, under CC-PDM-1.0.
- vibeace is “Vibe Ace” by Kevin MacLeod under CC-BY-4.0.
- trumpet is “solo trumpet 06” by Mihai Sorohan under CC-BY-4.0.
Catalog did not finish loading?
Confirm that your Mac has an internet connection for the first download. Rerun the script after the connection is available.
If a dependency or download problem continues, help me troubleshoot the librosa example catalog load.
Create the shifted noisy query
The test query comes from the middle of vibeace instead of its opening. A fixed NumPy seed makes the added Gaussian noise reproducible across runs.
- Insert the query and noise functions between load_catalog() and main() by pasting this code:
def choose_query(audio: object, offset_fraction: float) -> tuple[object, float]:
max_start = max(0, len(audio) - QUERY_SAMPLES)
start = int(max_start * offset_fraction)
return audio[start : start + QUERY_SAMPLES], start / SAMPLE_RATE
def add_white_noise(audio: object, snr_db: float, seed: int) -> object:
signal_power = float(np.mean(audio**2))
noise_power = signal_power / (10 ** (snr_db / 10))
rng = np.random.default_rng(seed)
noise = rng.normal(0.0, math.sqrt(noise_power), size=len(audio))
return audio + noise
How is the query controlled?
- choose_query() converts an offset fraction into a sample position before returning a two-second slice.
- add_white_noise() derives the noise power from the signal power and requested signal-to-noise ratio.
- np.random.default_rng() uses the supplied seed to reproduce the same noise samples.
- Save recognizer.py.
The file now contains choose_query() and add_white_noise() immediately above main().
Seeing a syntax warning in the query helpers?
Check that both functions sit outside load_catalog(). Each def line must begin at the left edge of the file.
For help comparing the function boundaries, help me fix the query helper placement.
The next version of main() selects the same shifted excerpt on every run. It also prints the calculated offset as a concrete check.
- Locate the existing main() function near the bottom of recognizer.py.
- Replace that function with the version below:
def main() -> None:
print("Loading the licensed reference catalog...")
catalog = load_catalog()
demo_song = "vibeace"
clean_query, true_offset = choose_query(catalog[demo_song], 0.45)
noisy_query = add_white_noise(clean_query, snr_db=15, seed=20261010)
print(f"Generated query offset: {true_offset:.2f}s")
What changed in main?
- demo_song selects vibeace as the known source recording.
- 0.45 places the query partway through the available reference audio.
- snr_db=15 adds a controlled level of noise.
- seed=20261010 keeps the noisy query deterministic.
- Save recognizer.py.
Before you rerun the script, make your prediction: the same offset or a changing offset.
- Generate the shifted noisy query by running:
python recognizer.py
What should this run show?
The fixed catalog duration and offset fraction produce a query offset of about 2.70s. Running the script again produces the same value.
You will see Generated query offset: 2.70s after the catalog loads. The experiment now has a repeatable shifted noisy query.
Offset missing from the output?
Confirm that you replaced only the main() function. Keep the final if __name__ == "__main__": block below it.
If the script stops before printing the offset, help me debug the deterministic query generation.
Compare the opening templates
The baseline uses cosine similarity to compare the query with the first two seconds of every recording. The acceptance threshold turns the best score into either a catalog prediction or a rejection.
- Insert the baseline functions between add_white_noise() and main() by pasting this code:
def cosine_similarity(left: object, right: object) -> float:
sample_count = min(len(left), len(right))
left = left[:sample_count]
right = right[:sample_count]
numerator = sum(float(a * b) for a, b in zip(left, right))
left_energy = math.sqrt(sum(float(value * value) for value in left))
right_energy = math.sqrt(sum(float(value * value) for value in right))
denominator = left_energy * right_energy
return numerator / denominator if denominator else 0.0
def build_baseline_templates(catalog: dict[str, object]) -> dict[str, object]:
return {song_id: audio[:QUERY_SAMPLES] for song_id, audio in catalog.items()}
def baseline_match(
query: object, templates: dict[str, object]
) -> tuple[str | None, float]:
best_song: str | None = None
best_score = float("-inf")
for song_id, template in templates.items():
score = cosine_similarity(query, template)
if score > best_score:
best_song = song_id
best_score = score
if best_score < BASELINE_ACCEPT_SCORE:
return None, best_score
return best_song, best_score
How does the baseline decide?
- cosine_similarity() normalizes the waveform comparison by the energy of both clips.
- build_baseline_templates() keeps only the opening two seconds from each catalog recording.
- baseline_match() scans every opening template and retains its highest score.
- BASELINE_ACCEPT_SCORE rejects the winner when that score remains below the required threshold.
- Save recognizer.py.
The file now contains the complete template-building and baseline-comparison path above main().
Baseline helpers showing indentation problems?
Keep each function at the left edge of the file. Preserve the extra indentation inside the loop and both conditional blocks.
For a focused comparison, help me fix the baseline helper indentation.
The final baseline path builds the opening templates before scoring the shifted query. It prints the selected label with the raw similarity score.
- Locate the current main() function near the bottom of recognizer.py.
- Replace that function with the completed baseline version below:
def main() -> None:
print("Loading the licensed reference catalog...")
catalog = load_catalog()
templates = build_baseline_templates(catalog)
demo_song = "vibeace"
clean_query, true_offset = choose_query(catalog[demo_song], 0.45)
noisy_query = add_white_noise(clean_query, snr_db=15, seed=20261010)
baseline_song, baseline_score = baseline_match(noisy_query, templates)
baseline_label = (
CATALOG_LABELS[baseline_song] if baseline_song is not None else "No match"
)
print(f"Baseline prediction: {baseline_label} | score={baseline_score:.3f}")
print(f"Generated query offset: {true_offset:.2f}s")
How does the experiment come together?
- templates contains only the beginning of each catalog track.
- noisy_query contains a shifted section from the middle of vibeace.
- baseline_match() scores that query against every opening template.
- baseline_label converts a rejected score into No match.
- Save recognizer.py.
Before you run the completed baseline, make your prediction: recognition or rejection.
- Test the opening-template matcher from the active VS Code terminal by running:
python recognizer.py
What does the command test?
The command rebuilds the same catalog and query before applying the strict template matcher. The printed label exposes whether raw opening samples can identify the shifted excerpt.
You will see Baseline prediction: No match with a score below 0.80. The correct recording is present, yet the shifted samples do not line up with its opening template.
That shortfall is intentional. You have exposed the offset-alignment failure that makes direct waveform templates unsuitable for this retrieval task.
What does this reveal about system design?
The baseline scans one template per catalog track, so its comparison work grows with the catalog. More machines can increase throughput without correcting the raw alignment problem.
An inverted index needs a representation that supports direct lookups from stable audio evidence. The failed baseline shows why the retrieval primitive matters before scaling begins.
Baseline did not print No match?
Confirm that build_baseline_templates() slices each recording from its beginning. Confirm that choose_query() uses the 0.45 offset fraction.
Check that BASELINE_ACCEPT_SCORE remains 0.80 and the noise seed remains 20261010.
If the result still differs, help me compare my baseline recognizer with the expected deterministic settings.
✔️ Awesome, I've got everything!
Great. Your saved baseline now rejects the shifted noisy vibeace query in a reproducible way.
ⓧ I'd like to double check the full code
Compare your complete recognizer.py with this baseline-stage file.
from __future__ import annotations
import math
from dataclasses import dataclass
import librosa
import numpy as np
SAMPLE_RATE = 11_025
CATALOG_SECONDS = 8.0
QUERY_SECONDS = 2.0
QUERY_SAMPLES = int(SAMPLE_RATE * QUERY_SECONDS)
BASELINE_ACCEPT_SCORE = 0.80
CATALOG_KEYS = ("brahms", "vibeace", "trumpet")
CATALOG_LABELS = {
"brahms": "Hungarian Dance number 5",
"vibeace": "Vibe Ace",
"trumpet": "solo trumpet 06",
}
def load_catalog() -> dict[str, object]:
catalog: dict[str, object] = {}
for key in CATALOG_KEYS:
path = librosa.example(key)
audio, _ = librosa.load(
path,
sr=SAMPLE_RATE,
mono=True,
duration=CATALOG_SECONDS,
)
catalog[key] = audio
return catalog
def choose_query(audio: object, offset_fraction: float) -> tuple[object, float]:
max_start = max(0, len(audio) - QUERY_SAMPLES)
start = int(max_start * offset_fraction)
return audio[start : start + QUERY_SAMPLES], start / SAMPLE_RATE
def add_white_noise(audio: object, snr_db: float, seed: int) -> object:
signal_power = float(np.mean(audio**2))
noise_power = signal_power / (10 ** (snr_db / 10))
rng = np.random.default_rng(seed)
noise = rng.normal(0.0, math.sqrt(noise_power), size=len(audio))
return audio + noise
def cosine_similarity(left: object, right: object) -> float:
sample_count = min(len(left), len(right))
left = left[:sample_count]
right = right[:sample_count]
numerator = sum(float(a * b) for a, b in zip(left, right))
left_energy = math.sqrt(sum(float(value * value) for value in left))
right_energy = math.sqrt(sum(float(value * value) for value in right))
denominator = left_energy * right_energy
return numerator / denominator if denominator else 0.0
def build_baseline_templates(catalog: dict[str, object]) -> dict[str, object]:
return {song_id: audio[:QUERY_SAMPLES] for song_id, audio in catalog.items()}
def baseline_match(
query: object, templates: dict[str, object]
) -> tuple[str | None, float]:
best_song: str | None = None
best_score = float("-inf")
for song_id, template in templates.items():
score = cosine_similarity(query, template)
if score > best_score:
best_song = song_id
best_score = score
if best_score < BASELINE_ACCEPT_SCORE:
return None, best_score
return best_song, best_score
def main() -> None:
print("Loading the licensed reference catalog...")
catalog = load_catalog()
templates = build_baseline_templates(catalog)
demo_song = "vibeace"
clean_query, true_offset = choose_query(catalog[demo_song], 0.45)
noisy_query = add_white_noise(clean_query, snr_db=15, seed=20261010)
baseline_song, baseline_score = baseline_match(noisy_query, templates)
baseline_label = (
CATALOG_LABELS[baseline_song] if baseline_song is not None else "No match"
)
print(f"Baseline prediction: {baseline_label} | score={baseline_score:.3f}")
print(f"Generated query offset: {true_offset:.2f}s")
if __name__ == "__main__":
main()
Your deterministic baseline now fails for the exact reason this experiment was designed to expose. Next, you will replace raw waveform alignment with sparse time-frequency landmarks.
Create the Constellation Map
Your shifted noisy query already defeats the waveform baseline. The recognizer now needs evidence that survives a different start time.
The failed comparison needs a sparse representation of the strongest time-frequency landmarks. A spectrogram provides that representation while preserving clues that can survive added noise.
In this step, get ready to:
- Convert the noisy query into a decibel spectrogram.
- Retain local spectral peaks above an amplitude threshold.
- Save a constellation map with cyan peak markers.
Transform audio into a spectrogram
A short-time Fourier transform splits audio into overlapping windows. Each window reveals which frequencies carry energy during that moment.
- In recognizer.py, find the section that starts with import librosa.
- Replace the section through CATALOG_SECONDS = 8.0 with this code:
import librosa
import matplotlib.pyplot as plt
import numpy as np
from scipy.ndimage import maximum_filter
SAMPLE_RATE = 11_025
N_FFT = 1_024
HOP_LENGTH = 256
CATALOG_SECONDS = 8.0
What do these additions control?
- Matplotlib renders the spectrogram as an image.
- SciPy supplies the two-dimensional local maximum filter used in the next substep.
- N_FFT sets the number of samples examined in each transform window.
- HOP_LENGTH sets how far the transform advances between windows.
- In recognizer.py, add spectrogram_db() directly below baseline_match() by pasting this code:
def spectrogram_db(audio: object) -> object:
magnitude = np.abs(
librosa.stft(
y=audio,
n_fft=N_FFT,
hop_length=HOP_LENGTH,
center=True,
)
)
return librosa.amplitude_to_db(magnitude, ref=np.max, top_db=80.0)
What does this code do?
- The transform converts the waveform into a time-frequency matrix.
- The absolute value keeps the magnitude of each frequency bin.
- The decibel conversion expresses each cell relative to the strongest magnitude.
- A larger FFT window improves frequency resolution. A shorter hop improves time resolution.
- More detailed windows also increase feature-extraction work for every query.
- Save recognizer.py.
- Confirm the new transform code preserves the working baseline by running:
python recognizer.py
What should you see?
You should still see Baseline prediction: No match. This regression check proves the new transform function loads without breaking the baseline.
Seeing an import or syntax error?
Check that your terminal still shows the active .venv environment. Confirm that each new import sits at the top of recognizer.py.
Compare the indentation inside spectrogram_db() with the snippet above.
Help me debug the spectrogram setup.
Keep the local spectral peaks
Most spectrogram cells contain weak or redundant energy. Local spectral peaks keep the strongest point within each time-frequency neighbourhood.
The amplitude threshold removes quieter cells before they become retrieval evidence. This produces the sparse constellation needed for compact fingerprints.
- In recognizer.py, find QUERY_SAMPLES = int(SAMPLE_RATE * QUERY_SECONDS).
- Add the peak controls directly below that line by pasting:
PEAK_NEIGHBORHOOD = (15, 7)
PEAK_THRESHOLD_DB = -35.0
How does peak pruning work?
PEAK_NEIGHBORHOOD defines the area that competes for one local maximum. PEAK_THRESHOLD_DB removes cells more than 35 decibels below the spectrogram reference.
A wider neighbourhood stores fewer peaks. A stricter threshold can remove useful evidence from noisy queries.
- Add find_peaks() directly below spectrogram_db() by pasting this code:
def find_peaks(db_spectrogram: object) -> list[tuple[int, int]]:
local_maximum = maximum_filter(
db_spectrogram,
size=PEAK_NEIGHBORHOOD,
mode="constant",
cval=-80.0,
)
peak_mask = (db_spectrogram == local_maximum) & (
db_spectrogram > PEAK_THRESHOLD_DB
)
coordinates = np.argwhere(peak_mask)
return sorted((int(time), int(frequency)) for frequency, time in coordinates)
What does this code do?
- maximum_filter assigns the strongest nearby value to each spectrogram cell.
- peak_mask keeps cells that equal their neighbourhood maximum.
- The threshold condition discards weak maxima.
- The final list stores each peak as time first. This ordering prepares the landmarks for time-offset matching.
- Save recognizer.py.
- Check that the peak extraction code parses correctly by running:
python recognizer.py
What should you see?
You should see Baseline prediction: No match again. The peak extractor now loads successfully alongside the original matcher.
Peak extraction code not running?
Confirm that PEAK_NEIGHBORHOOD contains two values inside parentheses. Check that find_peaks() sits outside spectrogram_db().
Help me fix the peak extractor.
Plot the constellation map
The spectrogram shows the complete time-frequency surface. The constellation overlay reveals how much of that surface survives peak pruning.
- Add save_constellation() directly below find_peaks() by pasting this first visual layer:
def save_constellation(audio: object, output_path: str = "constellation.png") -> None:
db_spectrogram = spectrogram_db(audio)
duration_seconds = len(audio) / SAMPLE_RATE
max_frequency = SAMPLE_RATE / 2
plt.figure(figsize=(10, 4), layout="tight")
plt.imshow(
db_spectrogram,
origin="lower",
aspect="auto",
extent=(0, duration_seconds, 0, max_frequency),
cmap="magma",
vmin=-80,
vmax=0,
)
plt.title("Noisy query spectrogram and retained landmark peaks")
plt.xlabel("Time in seconds")
plt.savefig(output_path, dpi=150, bbox_inches="tight")
What does this visual layer do?
- The horizontal axis represents the query duration in seconds.
- The vertical axis represents frequencies up to half the sample rate.
- The magma colour scale maps quieter cells to darker colours.
- The save operation writes the current figure to constellation.png.
- In recognizer.py, replace the existing main() function plus its final launch condition with this version:
def main() -> None:
print("Loading the licensed reference catalog...")
catalog = load_catalog()
templates = build_baseline_templates(catalog)
demo_song = "vibeace"
clean_query, true_offset = choose_query(catalog[demo_song], 0.45)
noisy_query = add_white_noise(clean_query, snr_db=15, seed=20261010)
baseline_song, baseline_score = baseline_match(noisy_query, templates)
save_constellation(noisy_query)
baseline_label = (
CATALOG_LABELS[baseline_song] if baseline_song is not None else "No match"
)
print(f"Baseline prediction: {baseline_label} | score={baseline_score:.3f}")
print("Saved constellation.png")
if __name__ == "__main__":
main()
How does the image enter the workflow?
main() passes the same deterministic noisy query to save_constellation(). The baseline output remains available for comparison.
The final message confirms that the image-writing path completed.
- Save recognizer.py.
Before you run this check, what visual pattern do you expect the noisy audio to create across time and frequency?
- Generate the first image layer by running:
python recognizer.py
What should you see?
The terminal should report the failed baseline plus Saved constellation.png. Open constellation.png from the VS Code Explorer sidebar to see the magma spectrogram.
The colour bands show the full energy surface. The next edit overlays the sparse landmarks retained for matching.
Image file missing?
Confirm that save_constellation(noisy_query) sits inside main() after the baseline comparison. Check the terminal for an earlier plotting error.
Help me generate the constellation image.
- In save_constellation(), find db_spectrogram = spectrogram_db(audio).
- Add this line directly below it:
peaks = find_peaks(db_spectrogram)
Why calculate peaks here?
The plot and the peak extractor now read the same decibel matrix. This keeps the heatmap aligned with every marker placed over it.
- In the same function, find the line that sets the plot title.
- Insert this marker layer directly above that title line:
if peaks:
peak_times = [time * HOP_LENGTH / SAMPLE_RATE for time, _ in peaks]
peak_frequencies = [
frequency * SAMPLE_RATE / N_FFT for _, frequency in peaks
]
plt.scatter(
peak_times,
peak_frequencies,
s=8,
c="cyan",
alpha=0.8,
edgecolors="none",
)
What does the overlay show?
- peak_times converts transform frames into seconds.
- peak_frequencies converts frequency bins into hertz.
- The cyan points mark the sparse cells retained as landmarks.
- Fewer landmarks reduce future storage work. Excessive pruning can reduce recall under noise.
- Save recognizer.py.
Before you run the final check, do you expect the retained points to cover every bright cell or form a sparse pattern?
- Regenerate the completed constellation map by running:
python recognizer.py
What should you see?
You should still see Baseline prediction: No match in the terminal. You should also see Saved constellation.png.
Open the regenerated image. You should see sparse cyan peaks over the magma spectrogram under the title Noisy query spectrogram and retained landmark peaks.
That is the key representation shift. Your recognizer now keeps compact landmarks instead of depending on raw sample alignment.
Cyan peaks not visible?
Confirm that peaks = find_peaks(db_spectrogram) appears near the top of save_constellation(). Check that the marker block sits before the title and save lines.
Make sure you reopened the regenerated image in VS Code. The editor can keep an older preview visible until you close the image tab.
Help me debug the missing peak markers.
✔️ Awesome, I've got everything!
Strong work. Your noisy query now becomes a sparse constellation map while the shifted waveform baseline still reports no match.
Double check that recognizer.py is saved. Keep constellation.png in the same project folder.
ⓧ I'd like to double check the full code
Compare your complete recognizer.py with this constellation-stage version.
from __future__ import annotations
import math
from dataclasses import dataclass
import librosa
import matplotlib.pyplot as plt
import numpy as np
from scipy.ndimage import maximum_filter
SAMPLE_RATE = 11_025
N_FFT = 1_024
HOP_LENGTH = 256
CATALOG_SECONDS = 8.0
QUERY_SECONDS = 2.0
QUERY_SAMPLES = int(SAMPLE_RATE * QUERY_SECONDS)
PEAK_NEIGHBORHOOD = (15, 7)
PEAK_THRESHOLD_DB = -35.0
BASELINE_ACCEPT_SCORE = 0.80
CATALOG_KEYS = ("brahms", "vibeace", "trumpet")
CATALOG_LABELS = {
"brahms": "Hungarian Dance number 5",
"vibeace": "Vibe Ace",
"trumpet": "solo trumpet 06",
}
def load_catalog() -> dict[str, object]:
catalog: dict[str, object] = {}
for key in CATALOG_KEYS:
path = librosa.example(key)
audio, _ = librosa.load(
path,
sr=SAMPLE_RATE,
mono=True,
duration=CATALOG_SECONDS,
)
catalog[key] = audio
return catalog
def choose_query(audio: object, offset_fraction: float) -> tuple[object, float]:
max_start = max(0, len(audio) - QUERY_SAMPLES)
start = int(max_start * offset_fraction)
return audio[start : start + QUERY_SAMPLES], start / SAMPLE_RATE
def add_white_noise(audio: object, snr_db: float, seed: int) -> object:
signal_power = float(np.mean(audio**2))
noise_power = signal_power / (10 ** (snr_db / 10))
rng = np.random.default_rng(seed)
noise = rng.normal(0.0, math.sqrt(noise_power), size=len(audio))
return audio + noise
def cosine_similarity(left: object, right: object) -> float:
sample_count = min(len(left), len(right))
left = left[:sample_count]
right = right[:sample_count]
numerator = sum(float(a * b) for a, b in zip(left, right))
left_energy = math.sqrt(sum(float(value * value) for value in left))
right_energy = math.sqrt(sum(float(value * value) for value in right))
denominator = left_energy * right_energy
return numerator / denominator if denominator else 0.0
def build_baseline_templates(catalog: dict[str, object]) -> dict[str, object]:
return {song_id: audio[:QUERY_SAMPLES] for song_id, audio in catalog.items()}
def baseline_match(
query: object, templates: dict[str, object]
) -> tuple[str | None, float]:
best_song: str | None = None
best_score = float("-inf")
for song_id, template in templates.items():
score = cosine_similarity(query, template)
if score > best_score:
best_song = song_id
best_score = score
if best_score < BASELINE_ACCEPT_SCORE:
return None, best_score
return best_song, best_score
def spectrogram_db(audio: object) -> object:
magnitude = np.abs(
librosa.stft(
y=audio,
n_fft=N_FFT,
hop_length=HOP_LENGTH,
center=True,
)
)
return librosa.amplitude_to_db(magnitude, ref=np.max, top_db=80.0)
def find_peaks(db_spectrogram: object) -> list[tuple[int, int]]:
local_maximum = maximum_filter(
db_spectrogram,
size=PEAK_NEIGHBORHOOD,
mode="constant",
cval=-80.0,
)
peak_mask = (db_spectrogram == local_maximum) & (
db_spectrogram > PEAK_THRESHOLD_DB
)
coordinates = np.argwhere(peak_mask)
return sorted((int(time), int(frequency)) for frequency, time in coordinates)
def save_constellation(audio: object, output_path: str = "constellation.png") -> None:
db_spectrogram = spectrogram_db(audio)
peaks = find_peaks(db_spectrogram)
duration_seconds = len(audio) / SAMPLE_RATE
max_frequency = SAMPLE_RATE / 2
plt.figure(figsize=(10, 4), layout="tight")
plt.imshow(
db_spectrogram,
origin="lower",
aspect="auto",
extent=(0, duration_seconds, 0, max_frequency),
cmap="magma",
vmin=-80,
vmax=0,
)
if peaks:
peak_times = [time * HOP_LENGTH / SAMPLE_RATE for time, _ in peaks]
peak_frequencies = [
frequency * SAMPLE_RATE / N_FFT for _, frequency in peaks
]
plt.scatter(
peak_times,
peak_frequencies,
s=8,
c="cyan",
alpha=0.8,
edgecolors="none",
)
plt.title("Noisy query spectrogram and retained landmark peaks")
plt.xlabel("Time in seconds")
plt.savefig(output_path, dpi=150, bbox_inches="tight")
def main() -> None:
print("Loading the licensed reference catalog...")
catalog = load_catalog()
templates = build_baseline_templates(catalog)
demo_song = "vibeace"
clean_query, true_offset = choose_query(catalog[demo_song], 0.45)
noisy_query = add_white_noise(clean_query, snr_db=15, seed=20261010)
baseline_song, baseline_score = baseline_match(noisy_query, templates)
save_constellation(noisy_query)
baseline_label = (
CATALOG_LABELS[baseline_song] if baseline_song is not None else "No match"
)
print(f"Baseline prediction: {baseline_label} | score={baseline_score:.3f}")
print("Saved constellation.png")
if __name__ == "__main__":
main()
Your file should contain the baseline functions plus the new spectrogram, peak extraction, and plotting functions. Save any corrections before running the script again.
Your recognizer now turns noisy audio into sparse time-frequency landmarks. Next, you will pair those landmarks into hashes so matching can jump directly to relevant catalog postings.
Index Fingerprints and Vote on Offsets
Your constellation map now preserves the strongest time-frequency peaks from a shifted noisy query. The baseline still misses the correct recording because it compares raw samples at one fixed alignment.
A single peak can appear in several recordings. You will turn nearby peaks into an audio fingerprint. An inverted index retrieves matching recordings by hash. Time-offset voting then rewards evidence that agrees on one alignment.
In this step, get ready to:
- Pair nearby spectral peaks into discriminative fingerprint hashes.
- Store each hash with its recording and reference time in an inverted index.
- Identify the recording by voting on matching time offsets.
Pair peaks into fingerprints
A peak pair captures two frequencies plus the time between them. This combination gives each hash more context than an isolated peak.
- In recognizer.py, find the import math line.
- Replace that line with this import group:
import math
from collections import Counter, defaultdict
from dataclasses import dataclass
Why add these imports?
- Counter finds the strongest song-and-offset vote cluster.
- defaultdict collects postings under each fingerprint hash.
- dataclass gives every match result one consistent structure.
- Find PEAK_THRESHOLD_DB = -35.0 near the top of recognizer.py.
- Add these fingerprint controls directly below it:
FAN_OUT = 5
TARGET_MIN_FRAMES = 2
TARGET_MAX_FRAMES = 30
FREQUENCY_QUANTIZATION = 2
What do these controls define?
- FAN_OUT limits each anchor peak to five target peaks.
- TARGET_MIN_FRAMES skips targets that are too close to the anchor.
- TARGET_MAX_FRAMES stops pair creation beyond the matching time window.
- FREQUENCY_QUANTIZATION groups nearby frequency bins into the same hash coordinate.
- Find the closing brace of CATALOG_LABELS.
- Add the fingerprint aliases plus the match result model below that dictionary:
HashKey = tuple[int, int, int]
Fingerprint = tuple[HashKey, int]
Posting = tuple[str, int]
@dataclass(frozen=True)
class Match:
song_id: str | None
offset_seconds: float | None
votes: int
confidence: float
matched_postings: int
How is matching evidence represented?
- HashKey stores the quantized anchor frequency, target frequency, and frame delta.
- Fingerprint combines a hash with its anchor time.
- Posting records a song ID plus the hash time in that recording.
- Match keeps the winner, estimated offset, vote count, confidence, and posting count together.
- Find the end of find_peaks() in recognizer.py.
- Add the complete fingerprint() function below it:
def fingerprint(audio: object) -> list[Fingerprint]:
peaks = find_peaks(spectrogram_db(audio))
hashes: list[Fingerprint] = []
for anchor_index, (anchor_time, anchor_frequency) in enumerate(peaks):
paired = 0
for target_time, target_frequency in peaks[anchor_index + 1 :]:
delta_time = target_time - anchor_time
if delta_time < TARGET_MIN_FRAMES:
continue
if delta_time > TARGET_MAX_FRAMES:
break
hash_key = (
anchor_frequency // FREQUENCY_QUANTIZATION,
target_frequency // FREQUENCY_QUANTIZATION,
delta_time,
)
hashes.append((hash_key, anchor_time))
paired += 1
if paired == FAN_OUT:
break
return hashes
What does this code do?
- find_peaks() supplies the sparse constellation points for the recording.
- delta_time measures the target peak's distance from the anchor in spectrogram frames.
- hash_key combines two quantized frequencies with their time delta.
- paired caps the work performed for each anchor peak.
- Save recognizer.py.
- Confirm the recognizer remains runnable by running:
python recognizer.py
What does this checkpoint prove?
You should still see the baseline reject the shifted query. You should also see constellation.png saved successfully.
This confirms the fingerprint definitions load without disrupting the working baseline or peak plot.
Seeing an error after adding fingerprints?
Check that the new imports sit with the existing imports. Confirm that the aliases and Match class appear after CATALOG_LABELS.
Check the indentation inside both loops in fingerprint(). The inner loop must remain nested under the anchor loop.
Ask for help with the fingerprint code.
Build the hash index
The index stores a posting list under every hash key. Query processing can retrieve those short lists directly instead of comparing the query with every full waveform.
- Add build_index() directly below fingerprint() in recognizer.py by pasting:
def build_index(
catalog: dict[str, object],
) -> tuple[dict[HashKey, list[Posting]], dict[str, int]]:
index: defaultdict[HashKey, list[Posting]] = defaultdict(list)
per_song_hashes: dict[str, int] = {}
for song_id, audio in catalog.items():
hashes = fingerprint(audio)
per_song_hashes[song_id] = len(hashes)
for hash_key, reference_time in hashes:
index[hash_key].append((song_id, reference_time))
return dict(index), per_song_hashes
How does the index work?
- index maps each fingerprint hash to its matching song and reference-time postings.
- per_song_hashes records how many fingerprints each catalog track contributes.
- dict(index) returns an ordinary dictionary for the lookup stage.
- Find the existing main() function near the bottom of recognizer.py.
- Replace that function plus its entry point with this index checkpoint:
def main() -> None:
print("Loading the licensed reference catalog...")
catalog = load_catalog()
templates = build_baseline_templates(catalog)
index, per_song_hashes = build_index(catalog)
demo_song = "vibeace"
clean_query, true_offset = choose_query(catalog[demo_song], 0.45)
noisy_query = add_white_noise(clean_query, snr_db=15, seed=20261010)
baseline_song, baseline_score = baseline_match(noisy_query, templates)
save_constellation(noisy_query)
baseline_label = (
CATALOG_LABELS[baseline_song] if baseline_song is not None else "No match"
)
posting_count = sum(len(postings) for postings in index.values())
print(f"Baseline prediction: {baseline_label} | score={baseline_score:.3f}")
print(f"Generated query offset: {true_offset:.2f}s")
print(f"Unique hash keys: {len(index)} | postings: {posting_count}")
print(f"Hashes per song: {per_song_hashes}")
print("Saved constellation.png")
if __name__ == "__main__":
main()
What does this checkpoint add?
main() now fingerprints all three catalog recordings. It calculates the number of unique keys plus the total number of postings.
Those counts describe the index's shape without treating Python object counts as production storage bytes.
- Save recognizer.py.
- Build the index and inspect its shape by running:
python recognizer.py
What should you see?
You should see Unique hash keys: followed by a nonzero key count. The same line should report a nonzero posting count.
You should also see Hashes per song: with entries for brahms, vibeace, and trumpet.
Seeing zero keys or no index summary?
Confirm that build_index(catalog) appears after the template setup in main(). Check that build_index() calls fingerprint(audio) for every catalog entry.
Check that index[hash_key].append((song_id, reference_time)) remains inside the inner loop.
Ask for help with the index output.
Vote on matching offsets
A matching hash only identifies possible evidence. A real excerpt produces many postings that agree on the same song plus the same difference between reference time and query time.
- Add vote_matches() directly below build_index() in recognizer.py by pasting:
def vote_matches(
query_hashes: list[Fingerprint], index: dict[HashKey, list[Posting]]
) -> Match:
votes: Counter[tuple[str, int]] = Counter()
matched_postings = 0
for hash_key, query_time in query_hashes:
for song_id, reference_time in index.get(hash_key, []):
votes[(song_id, reference_time - query_time)] += 1
matched_postings += 1
if not votes:
return Match(None, None, 0, 0.0, 0)
(song_id, offset_frames), best_votes = votes.most_common(1)[0]
total_votes = sum(votes.values())
confidence = best_votes / total_votes
offset_seconds = offset_frames * HOP_LENGTH / SAMPLE_RATE
return Match(
song_id,
offset_seconds,
best_votes,
confidence,
matched_postings,
)
How does offset voting find the winner?
- index.get(hash_key, []) retrieves only the postings that share a query hash.
- reference_time - query_time estimates where the query begins inside the reference recording.
- votes.most_common(1) selects the strongest coherent song-and-offset cluster.
- confidence measures the winning cluster's share of all offset votes.
- Add the raw-audio wrapper plus the output formatter below vote_matches() by pasting:
def raw_match_audio(query: object, index: dict[HashKey, list[Posting]]) -> Match:
return vote_matches(fingerprint(query), index)
def format_match(match: Match) -> str:
if match.song_id is None:
return "No raw hash match"
return (
f"{CATALOG_LABELS[match.song_id]} | "
f"offset={match.offset_seconds:.2f}s | "
f"votes={match.votes} | confidence={match.confidence:.3f}"
)
What do these helpers do?
raw_match_audio() fingerprints a query before passing its hashes to the vote aggregator.
format_match() turns the result into a readable label, offset, vote count, and confidence score.
- Replace the temporary main() checkpoint plus its entry point with the final version:
def main() -> None:
print("Loading the licensed reference catalog...")
catalog = load_catalog()
templates = build_baseline_templates(catalog)
index, per_song_hashes = build_index(catalog)
demo_song = "vibeace"
clean_query, true_offset = choose_query(catalog[demo_song], 0.45)
noisy_query = add_white_noise(clean_query, snr_db=15, seed=20261010)
baseline_song, baseline_score = baseline_match(noisy_query, templates)
fingerprint_match = raw_match_audio(noisy_query, index)
save_constellation(noisy_query)
baseline_label = (
CATALOG_LABELS[baseline_song] if baseline_song is not None else "No match"
)
posting_count = sum(len(postings) for postings in index.values())
print(f"Baseline prediction: {baseline_label} | score={baseline_score:.3f}")
print(f"Fingerprint prediction: {format_match(fingerprint_match)}")
print(f"Generated query offset: {true_offset:.2f}s")
print(f"Unique hash keys: {len(index)} | postings: {posting_count}")
print(f"Hashes per song: {per_song_hashes}")
print("Saved constellation.png")
if __name__ == "__main__":
main()
How does the complete demo flow work?
The same shifted noisy query now goes through both retrieval methods. The baseline compares it with fixed waveform templates.
The fingerprint path extracts hashes, retrieves postings, and selects the strongest aligned vote cluster.
- Save recognizer.py.
Before you run the completed recognizer, which method do you expect to identify the shifted noisy excerpt?
- Run the completed recognizer with:
python recognizer.py
What proves the matcher works?
You should see Baseline prediction: No match. You should then see Fingerprint prediction: Vibe Ace with a positive vote count.
Compare the fingerprint offset with Generated query offset:. The two values should be close because the winning hashes agree on the query's position inside the reference.
That is the retrieval breakthrough. Your recognizer now identifies the shifted noisy recording while the raw waveform baseline rejects it.
Fingerprint result missing or incorrect?
Confirm that raw_match_audio() calls fingerprint(query). Confirm that it passes the result into vote_matches().
Check the vote key carefully. It must use reference_time - query_time so matching hashes align around the same offset.
Ask for help with the final match.
✔️ Awesome, I've got everything!
Your recognizer now indexes peak-pair fingerprints and identifies the shifted noisy vibeace query through coherent offset votes. Double-check that recognizer.py is saved.
ⓧ I'd like to double check the full code
Compare your complete recognizer.py with this reference:
from __future__ import annotations
import math
from collections import Counter, defaultdict
from dataclasses import dataclass
import librosa
import matplotlib.pyplot as plt
import numpy as np
from scipy.ndimage import maximum_filter
SAMPLE_RATE = 11_025
N_FFT = 1_024
HOP_LENGTH = 256
CATALOG_SECONDS = 8.0
QUERY_SECONDS = 2.0
QUERY_SAMPLES = int(SAMPLE_RATE * QUERY_SECONDS)
PEAK_NEIGHBORHOOD = (15, 7)
PEAK_THRESHOLD_DB = -35.0
FAN_OUT = 5
TARGET_MIN_FRAMES = 2
TARGET_MAX_FRAMES = 30
FREQUENCY_QUANTIZATION = 2
BASELINE_ACCEPT_SCORE = 0.80
CATALOG_KEYS = ("brahms", "vibeace", "trumpet")
CATALOG_LABELS = {
"brahms": "Hungarian Dance number 5",
"vibeace": "Vibe Ace",
"trumpet": "solo trumpet 06",
}
HashKey = tuple[int, int, int]
Fingerprint = tuple[HashKey, int]
Posting = tuple[str, int]
@dataclass(frozen=True)
class Match:
song_id: str | None
offset_seconds: float | None
votes: int
confidence: float
matched_postings: int
def load_catalog() -> dict[str, object]:
catalog: dict[str, object] = {}
for key in CATALOG_KEYS:
path = librosa.example(key)
audio, _ = librosa.load(
path,
sr=SAMPLE_RATE,
mono=True,
duration=CATALOG_SECONDS,
)
catalog[key] = audio
return catalog
def choose_query(audio: object, offset_fraction: float) -> tuple[object, float]:
max_start = max(0, len(audio) - QUERY_SAMPLES)
start = int(max_start * offset_fraction)
return audio[start : start + QUERY_SAMPLES], start / SAMPLE_RATE
def add_white_noise(audio: object, snr_db: float, seed: int) -> object:
signal_power = float(np.mean(audio**2))
noise_power = signal_power / (10 ** (snr_db / 10))
rng = np.random.default_rng(seed)
noise = rng.normal(0.0, math.sqrt(noise_power), size=len(audio))
return audio + noise
def cosine_similarity(left: object, right: object) -> float:
sample_count = min(len(left), len(right))
left = left[:sample_count]
right = right[:sample_count]
numerator = sum(float(a * b) for a, b in zip(left, right))
left_energy = math.sqrt(sum(float(value * value) for value in left))
right_energy = math.sqrt(sum(float(value * value) for value in right))
denominator = left_energy * right_energy
return numerator / denominator if denominator else 0.0
def build_baseline_templates(catalog: dict[str, object]) -> dict[str, object]:
return {song_id: audio[:QUERY_SAMPLES] for song_id, audio in catalog.items()}
def baseline_match(
query: object, templates: dict[str, object]
) -> tuple[str | None, float]:
best_song: str | None = None
best_score = float("-inf")
for song_id, template in templates.items():
score = cosine_similarity(query, template)
if score > best_score:
best_song = song_id
best_score = score
if best_score < BASELINE_ACCEPT_SCORE:
return None, best_score
return best_song, best_score
def spectrogram_db(audio: object) -> object:
magnitude = np.abs(
librosa.stft(
y=audio,
n_fft=N_FFT,
hop_length=HOP_LENGTH,
center=True,
)
)
return librosa.amplitude_to_db(magnitude, ref=np.max, top_db=80.0)
def find_peaks(db_spectrogram: object) -> list[tuple[int, int]]:
local_maximum = maximum_filter(
db_spectrogram,
size=PEAK_NEIGHBORHOOD,
mode="constant",
cval=-80.0,
)
peak_mask = (db_spectrogram == local_maximum) & (
db_spectrogram > PEAK_THRESHOLD_DB
)
coordinates = np.argwhere(peak_mask)
return sorted((int(time), int(frequency)) for frequency, time in coordinates)
def fingerprint(audio: object) -> list[Fingerprint]:
peaks = find_peaks(spectrogram_db(audio))
hashes: list[Fingerprint] = []
for anchor_index, (anchor_time, anchor_frequency) in enumerate(peaks):
paired = 0
for target_time, target_frequency in peaks[anchor_index + 1 :]:
delta_time = target_time - anchor_time
if delta_time < TARGET_MIN_FRAMES:
continue
if delta_time > TARGET_MAX_FRAMES:
break
hash_key = (
anchor_frequency // FREQUENCY_QUANTIZATION,
target_frequency // FREQUENCY_QUANTIZATION,
delta_time,
)
hashes.append((hash_key, anchor_time))
paired += 1
if paired == FAN_OUT:
break
return hashes
def build_index(
catalog: dict[str, object],
) -> tuple[dict[HashKey, list[Posting]], dict[str, int]]:
index: defaultdict[HashKey, list[Posting]] = defaultdict(list)
per_song_hashes: dict[str, int] = {}
for song_id, audio in catalog.items():
hashes = fingerprint(audio)
per_song_hashes[song_id] = len(hashes)
for hash_key, reference_time in hashes:
index[hash_key].append((song_id, reference_time))
return dict(index), per_song_hashes
def vote_matches(
query_hashes: list[Fingerprint], index: dict[HashKey, list[Posting]]
) -> Match:
votes: Counter[tuple[str, int]] = Counter()
matched_postings = 0
for hash_key, query_time in query_hashes:
for song_id, reference_time in index.get(hash_key, []):
votes[(song_id, reference_time - query_time)] += 1
matched_postings += 1
if not votes:
return Match(None, None, 0, 0.0, 0)
(song_id, offset_frames), best_votes = votes.most_common(1)[0]
total_votes = sum(votes.values())
confidence = best_votes / total_votes
offset_seconds = offset_frames * HOP_LENGTH / SAMPLE_RATE
return Match(
song_id,
offset_seconds,
best_votes,
confidence,
matched_postings,
)
def raw_match_audio(query: object, index: dict[HashKey, list[Posting]]) -> Match:
return vote_matches(fingerprint(query), index)
def save_constellation(audio: object, output_path: str = "constellation.png") -> None:
db_spectrogram = spectrogram_db(audio)
peaks = find_peaks(db_spectrogram)
duration_seconds = len(audio) / SAMPLE_RATE
max_frequency = SAMPLE_RATE / 2
plt.figure(figsize=(10, 4), layout="tight")
plt.imshow(
db_spectrogram,
origin="lower",
aspect="auto",
extent=(0, duration_seconds, 0, max_frequency),
cmap="magma",
vmin=-80,
vmax=0,
)
if peaks:
peak_times = [time * HOP_LENGTH / SAMPLE_RATE for time, _ in peaks]
peak_frequencies = [
frequency * SAMPLE_RATE / N_FFT for _, frequency in peaks
]
plt.scatter(
peak_times,
peak_frequencies,
s=8,
c="cyan",
alpha=0.8,
edgecolors="none",
)
plt.title("Noisy query spectrogram and retained landmark peaks")
plt.xlabel("Time in seconds")
plt.savefig(output_path, dpi=150, bbox_inches="tight")
def format_match(match: Match) -> str:
if match.song_id is None:
return "No raw hash match"
return (
f"{CATALOG_LABELS[match.song_id]} | "
f"offset={match.offset_seconds:.2f}s | "
f"votes={match.votes} | confidence={match.confidence:.3f}"
)
def main() -> None:
print("Loading the licensed reference catalog...")
catalog = load_catalog()
templates = build_baseline_templates(catalog)
index, per_song_hashes = build_index(catalog)
demo_song = "vibeace"
clean_query, true_offset = choose_query(catalog[demo_song], 0.45)
noisy_query = add_white_noise(clean_query, snr_db=15, seed=20261010)
baseline_song, baseline_score = baseline_match(noisy_query, templates)
fingerprint_match = raw_match_audio(noisy_query, index)
save_constellation(noisy_query)
baseline_label = (
CATALOG_LABELS[baseline_song] if baseline_song is not None else "No match"
)
posting_count = sum(len(postings) for postings in index.values())
print(f"Baseline prediction: {baseline_label} | score={baseline_score:.3f}")
print(f"Fingerprint prediction: {format_match(fingerprint_match)}")
print(f"Generated query offset: {true_offset:.2f}s")
print(f"Unique hash keys: {len(index)} | postings: {posting_count}")
print(f"Hashes per song: {per_song_hashes}")
print("Saved constellation.png")
if __name__ == "__main__":
main()
This reference contains the complete core recognizer for the current project state.
Your mini-Shazam now retrieves sparse fingerprint postings and reconstructs the correct recording offset. Next, you will benchmark accuracy, false positives, feature cost, and lookup latency.
Benchmark and Design for Scale
Your recognizer now finds a shifted noisy excerpt through an inverted index and coherent offset votes. The next question is whether that result holds across every catalog track under controlled degradation.
A credible recognizer measures accuracy, false positives, feature cost, and lookup latency separately. You will build a deterministic benchmark, protect the core behavior with tests, and document how the local design scales into a production service.
In this step, get ready to:
- Benchmark both matching methods across fixed offsets and noise levels.
- Test noisy recognition, offset accuracy, and index contents.
- Document the production architecture and its scaling trade-offs.
Benchmark recognition and retrieval costs
One successful demo cannot reveal how the matcher behaves across the full catalog. A deterministic evaluation matrix gives both methods the same offsets, noise levels, and repeatable random inputs.
Why separate the timings?
Fingerprint extraction is stateless CPU work. Index lookup retrieves postings and aggregates votes.
Timing them separately shows whether future optimization belongs in fingerprint workers or the retrieval tier.
- Create an empty evaluate.py file beside recognizer.py using the file icon with a plus sign in the VS Code sidebar.
- Add the benchmark imports, evaluation values, and percentage helper by pasting this code into evaluate.py:
from statistics import mean
from time import perf_counter
import numpy as np
from recognizer import (
QUERY_SAMPLES,
add_white_noise,
baseline_match,
build_baseline_templates,
build_index,
choose_query,
fingerprint,
load_catalog,
vote_matches,
)
OFFSET_FRACTIONS = (0.10, 0.45, 0.80)
SNR_LEVELS_DB = (30, 15, 5)
def percent(part: int, whole: int) -> float:
return 100.0 * part / whole if whole else 0.0
What does this code prepare?
- The imports reuse the catalog, baseline, fingerprint, and voting functions you already tested.
- The three offset fractions sample queries near the beginning, middle, and end of each recording.
- The three signal-to-noise values apply increasingly difficult degradation.
- The percent() helper converts result counts into readable rates.
- Save evaluate.py.
You should see the imports at the top of the file and percent() at the bottom.
Seeing unresolved imports?
Confirm that evaluate.py sits beside recognizer.py. Confirm that VS Code still uses the active .venv interpreter.
Ask for help with the import locations if the warning remains: help me diagnose unresolved imports in evaluate.py
- Add the benchmark setup below percent() by pasting this code:
def main() -> None:
catalog = load_catalog()
templates = build_baseline_templates(catalog)
index, per_song_hashes = build_index(catalog)
baseline_correct = 0
fingerprint_correct = 0
total_known = 0
baseline_latencies = []
feature_latencies = []
lookup_latencies = []
print("song,offset_seconds,snr_db,baseline,fingerprint,votes,confidence")
if __name__ == "__main__":
main()
What does this setup do?
- The catalog feeds both the raw waveform templates and the fingerprint index.
- The correctness counters track known-query results for each matcher.
- The latency lists keep comparison, extraction, and lookup measurements separate.
- The header defines the evidence printed for every benchmark case.
- Save evaluate.py.
- Run the starter benchmark in the existing terminal with this command:
python evaluate.py
What should you see?
You should see a comma-separated header beginning with song. This proves the benchmark can load the catalog and build the index before evaluating cases.
Benchmark does not reach the header?
Confirm that the .venv remains active. Check that recognizer.py is in the same folder as evaluate.py.
Use this prompt if the traceback points into catalog or index creation: help me debug evaluate.py before its benchmark header prints
- Insert the known-query evaluation loop immediately above the if __name__ == "__main__": line:
for song_position, (song_id, audio) in enumerate(catalog.items()):
for offset_position, offset_fraction in enumerate(OFFSET_FRACTIONS):
clean_query, offset_seconds = choose_query(audio, offset_fraction)
for noise_position, snr_db in enumerate(SNR_LEVELS_DB):
seed = 1000 + song_position * 100 + offset_position * 10 + noise_position
query = add_white_noise(clean_query, snr_db=snr_db, seed=seed)
start = perf_counter()
baseline_song, _ = baseline_match(query, templates)
baseline_latencies.append((perf_counter() - start) * 1000)
start = perf_counter()
query_hashes = fingerprint(query)
feature_latencies.append((perf_counter() - start) * 1000)
start = perf_counter()
fingerprint_result = vote_matches(query_hashes, index)
lookup_latencies.append((perf_counter() - start) * 1000)
total_known += 1
baseline_correct += int(baseline_song == song_id)
fingerprint_correct += int(fingerprint_result.song_id == song_id)
print(
f"{song_id},{offset_seconds:.2f},{snr_db},"
f"{baseline_song},{fingerprint_result.song_id},"
f"{fingerprint_result.votes},"
f"{fingerprint_result.confidence:.3f}"
)
How does the evaluation matrix work?
- The nested loops cover every catalog song at three relative offsets and three noise levels.
- Each position contributes to a fixed seed, so the generated noise stays reproducible.
- The three timers isolate baseline comparison, fingerprint extraction, and index lookup.
- Each printed row records both predictions with the fingerprint vote evidence.
- Save evaluate.py.
Before you run the expanded benchmark, how many result rows do you expect from three songs, three offsets, and three noise levels? Take a moment to decide.
- Run the expanded benchmark with this command:
python evaluate.py
What should the matrix show?
You should see 27 data rows beneath the header. Each song appears at every offset and noise level.
The baseline column often contains None for shifted excerpts. The fingerprint column should keep identifying catalog songs from aligned hash evidence.
Seeing fewer rows or a loop error?
Check that every nested block keeps four additional spaces of indentation. Make sure the loop sits inside main() and above the entry point.
Ask for an indentation review if Python stops before all cases run: help me inspect the nested loops in evaluate.py
- Insert the deterministic noise probes and first summary values above the if __name__ == "__main__": line:
baseline_false_positives = 0
fingerprint_false_positives = 0
probe_count = 3
for seed in (9001, 9002, 9003):
rng = np.random.default_rng(seed)
noise_query = rng.normal(0.0, 0.1, size=QUERY_SAMPLES)
baseline_song, _ = baseline_match(noise_query, templates)
fingerprint_result = vote_matches(fingerprint(noise_query), index)
baseline_false_positives += int(baseline_song is not None)
fingerprint_false_positives += int(fingerprint_result.song_id is not None)
raw_sample_count = sum(len(audio) for audio in catalog.values())
posting_count = sum(len(postings) for postings in index.values())
print("\nSummary")
print(
f"Baseline known accuracy: "
f"{percent(baseline_correct, total_known):.1f}%"
)
print(
f"Fingerprint known accuracy: "
f"{percent(fingerprint_correct, total_known):.1f}%"
)
What do the probes measure?
- Three fixed seeds create reproducible Gaussian-noise queries with no catalog recording behind them.
- Any accepted song from those probes counts as a false positive.
- Raw samples and posting counts describe the storage shape without treating Python objects as production byte estimates.
- The first summary lines compare known-query accuracy for both retrieval methods.
- Save evaluate.py.
- Run the benchmark again with this command:
python evaluate.py
What should the summary show?
You should see a Summary section after the 27 cases. Its first two lines report known-query accuracy for the baseline and fingerprint matcher.
Summary missing after the matrix?
Make sure the probe block remains inside main(). It should align with the outer catalog loop instead of nesting inside it.
Use this prompt to inspect the block position: help me place the summary block correctly in evaluate.py
- Complete the summary immediately after the fingerprint accuracy output by adding this code:
print(
f"Baseline false-positive rate on noise: "
f"{percent(baseline_false_positives, probe_count):.1f}%"
)
print(
f"Fingerprint false-positive rate on noise: "
f"{percent(fingerprint_false_positives, probe_count):.1f}%"
)
print(f"Mean baseline comparison latency: {mean(baseline_latencies):.3f} ms")
print(f"Mean fingerprint extraction time: {mean(feature_latencies):.3f} ms")
print(f"Mean inverted-index lookup latency: {mean(lookup_latencies):.3f} ms")
print(f"Raw catalog samples: {raw_sample_count}")
print(f"Unique hash keys: {len(index)}")
print(f"Index postings: {posting_count}")
print(f"Hashes per song: {per_song_hashes}")
How should you read these measurements?
- False-positive rates reveal how often unstructured noise produces a candidate.
- Mean extraction time captures the CPU cost before retrieval begins.
- Mean lookup latency captures posting-list access and vote aggregation.
- Key, posting, and per-song hash counts expose index fan-out and storage shape.
- Save evaluate.py.
Before the complete benchmark runs, predict which measured operation will take longer on one Mac: fingerprint extraction or in-memory index lookup. Keep your prediction in mind.
- Run the complete benchmark with this command:
python evaluate.py
What should the complete benchmark report?
You should see summaries for both matching methods, both false-positive rates, three separate timing measurements, and three index-shape measures.
Your exact timings can vary with the machine. The separate labels preserve the comparison that matters.
Seeing missing or empty measurements?
Check that each latency append stays inside the innermost known-query loop. Confirm that the final print statements remain inside main().
Ask for help with the measurement flow if a list is empty: help me debug the benchmark measurements in evaluate.py
✔️ Awesome, I've got everything!
Your deterministic benchmark now compares correctness, false positives, latency, and index shape. Double check that evaluate.py is saved.
ⓧ I'd like to double check the full code
Compare your saved evaluate.py with this complete version.
from statistics import mean
from time import perf_counter
import numpy as np
from recognizer import (
QUERY_SAMPLES,
add_white_noise,
baseline_match,
build_baseline_templates,
build_index,
choose_query,
fingerprint,
load_catalog,
vote_matches,
)
OFFSET_FRACTIONS = (0.10, 0.45, 0.80)
SNR_LEVELS_DB = (30, 15, 5)
def percent(part: int, whole: int) -> float:
return 100.0 * part / whole if whole else 0.0
def main() -> None:
catalog = load_catalog()
templates = build_baseline_templates(catalog)
index, per_song_hashes = build_index(catalog)
baseline_correct = 0
fingerprint_correct = 0
total_known = 0
baseline_latencies = []
feature_latencies = []
lookup_latencies = []
print("song,offset_seconds,snr_db,baseline,fingerprint,votes,confidence")
for song_position, (song_id, audio) in enumerate(catalog.items()):
for offset_position, offset_fraction in enumerate(OFFSET_FRACTIONS):
clean_query, offset_seconds = choose_query(audio, offset_fraction)
for noise_position, snr_db in enumerate(SNR_LEVELS_DB):
seed = 1000 + song_position * 100 + offset_position * 10 + noise_position
query = add_white_noise(clean_query, snr_db=snr_db, seed=seed)
start = perf_counter()
baseline_song, _ = baseline_match(query, templates)
baseline_latencies.append((perf_counter() - start) * 1000)
start = perf_counter()
query_hashes = fingerprint(query)
feature_latencies.append((perf_counter() - start) * 1000)
start = perf_counter()
fingerprint_result = vote_matches(query_hashes, index)
lookup_latencies.append((perf_counter() - start) * 1000)
total_known += 1
baseline_correct += int(baseline_song == song_id)
fingerprint_correct += int(fingerprint_result.song_id == song_id)
print(
f"{song_id},{offset_seconds:.2f},{snr_db},"
f"{baseline_song},{fingerprint_result.song_id},"
f"{fingerprint_result.votes},"
f"{fingerprint_result.confidence:.3f}"
)
baseline_false_positives = 0
fingerprint_false_positives = 0
probe_count = 3
for seed in (9001, 9002, 9003):
rng = np.random.default_rng(seed)
noise_query = rng.normal(0.0, 0.1, size=QUERY_SAMPLES)
baseline_song, _ = baseline_match(noise_query, templates)
fingerprint_result = vote_matches(fingerprint(noise_query), index)
baseline_false_positives += int(baseline_song is not None)
fingerprint_false_positives += int(fingerprint_result.song_id is not None)
raw_sample_count = sum(len(audio) for audio in catalog.values())
posting_count = sum(len(postings) for postings in index.values())
print("\nSummary")
print(
f"Baseline known accuracy: "
f"{percent(baseline_correct, total_known):.1f}%"
)
print(
f"Fingerprint known accuracy: "
f"{percent(fingerprint_correct, total_known):.1f}%"
)
print(
f"Baseline false-positive rate on noise: "
f"{percent(baseline_false_positives, probe_count):.1f}%"
)
print(
f"Fingerprint false-positive rate on noise: "
f"{percent(fingerprint_false_positives, probe_count):.1f}%"
)
print(f"Mean baseline comparison latency: {mean(baseline_latencies):.3f} ms")
print(f"Mean fingerprint extraction time: {mean(feature_latencies):.3f} ms")
print(f"Mean inverted-index lookup latency: {mean(lookup_latencies):.3f} ms")
print(f"Raw catalog samples: {raw_sample_count}")
print(f"Unique hash keys: {len(index)}")
print(f"Index postings: {posting_count}")
print(f"Hashes per song: {per_song_hashes}")
if __name__ == "__main__":
main()
Protect the recognizer with tests
Benchmark rows reveal broad behavior, while automated tests protect specific guarantees. The core suite checks identity, aligned votes, offset tolerance, and nonempty index postings.
- Create an empty test_recognizer.py file beside recognizer.py using the file icon with a plus sign in the VS Code sidebar.
- Add the shared test setup and known noisy-query test by pasting this code into test_recognizer.py:
import unittest
from recognizer import (
add_white_noise,
build_index,
choose_query,
load_catalog,
raw_match_audio,
)
class RecognizerTests(unittest.TestCase):
@classmethod
def setUpClass(cls) -> None:
cls.catalog = load_catalog()
cls.index, _ = build_index(cls.catalog)
def test_known_noisy_offset_clip(self) -> None:
query, expected_offset = choose_query(self.catalog["vibeace"], 0.45)
query = add_white_noise(query, snr_db=15, seed=42)
result = raw_match_audio(query, self.index)
self.assertEqual(result.song_id, "vibeace")
self.assertGreater(result.votes, 0)
self.assertAlmostEqual(result.offset_seconds, expected_offset, delta=0.30)
What does this test prove?
- The class setup builds the catalog and index once for the suite.
- The query uses the known shifted vibeace excerpt with deterministic noise.
- The assertions require the correct song, nonzero aligned votes, and an offset within 0.30 seconds.
- Save test_recognizer.py.
- Run the first core test with this command:
python -m unittest -v
What should the first test report?
You should see test_known_noisy_offset_clip followed by a passing status. The final suite status should be OK.
Known noisy clip test failing?
Compare the seed, noise level, offset fraction, and expected song with the snippet. Confirm that recognizer.py still contains the complete fingerprint matcher from the previous step.
Use the failing assertion in this prompt: help me diagnose the noisy offset recognition test
- Complete the test file below the first test by adding this code:
def test_index_contains_postings(self) -> None:
self.assertGreater(len(self.index), 0)
self.assertGreater(sum(len(values) for values in self.index.values()), 0)
if __name__ == "__main__":
unittest.main()
What does the index test protect?
The test requires at least one unique fingerprint key. It also requires at least one posting across all keys.
Together, these checks catch an empty index before recognition results become misleading.
- Save test_recognizer.py.
Before you rerun the suite, do you expect both tests to reuse one catalog and index build? Take a moment to decide.
- Run the complete core test suite with this command:
python -m unittest -v
What should the complete suite show?
You should see both named tests pass. The suite should finish with Ran 2 tests and OK.
Index test failing?
Confirm that setUpClass() assigns both cls.catalog and cls.index. Check that the second test remains indented inside RecognizerTests.
Ask for help with the failing suite structure: help me debug test_recognizer.py
✔️ Awesome, I've got everything!
Both core tests now protect the recognizer's known-query behavior and index contents.
ⓧ I'd like to double check the full code
Compare your saved test_recognizer.py with this complete version.
import unittest
from recognizer import (
add_white_noise,
build_index,
choose_query,
load_catalog,
raw_match_audio,
)
class RecognizerTests(unittest.TestCase):
@classmethod
def setUpClass(cls) -> None:
cls.catalog = load_catalog()
cls.index, _ = build_index(cls.catalog)
def test_known_noisy_offset_clip(self) -> None:
query, expected_offset = choose_query(self.catalog["vibeace"], 0.45)
query = add_white_noise(query, snr_db=15, seed=42)
result = raw_match_audio(query, self.index)
self.assertEqual(result.song_id, "vibeace")
self.assertGreater(result.votes, 0)
self.assertAlmostEqual(result.offset_seconds, expected_offset, delta=0.30)
def test_index_contains_postings(self) -> None:
self.assertGreater(len(self.index), 0)
self.assertGreater(sum(len(values) for values in self.index.values()), 0)
if __name__ == "__main__":
unittest.main()
Document the production design
The local functions already mirror production responsibilities. A clear README connects those responsibilities to offline ingestion, online query processing, index shards, vote aggregation, and metadata retrieval.
- Create an empty README.md file beside recognizer.py using the file icon with a plus sign in the VS Code sidebar.
- Add the project overview and reproducible setup by pasting this content into README.md:
# Mini-Shazam
A small, explainable recording-identification system built to connect audio fingerprints with production system-design concepts.
## What it demonstrates
- A direct waveform baseline fails when a recording starts at a different offset.
- Sparse spectrogram peaks preserve useful evidence under added noise.
- Peak pairs create more discriminative hashes than individual peaks.
- An inverted index maps each hash directly to candidate song and time postings.
- Time-offset voting distinguishes a coherent match from random collisions.
- Deterministic evaluation separates feature-extraction cost from index lookup latency.
This is recording identification, not melody recognition. It expects a degraded excerpt of the same underlying recording and is not designed to identify humming, cover versions, or changed pitch and tempo.
## Setup on macOS
Python 3.12 or newer is required by the pinned scientific packages.
```bash
python -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt
```
Run the demo, benchmark, and tests:
```bash
python recognizer.py
python evaluate.py
python -m unittest -v
```
The first run downloads and caches the documented librosa example recordings. The demo creates `constellation.png`.
What does the opening document?
The overview defines the recognizer's observable behavior and its boundaries. The setup section preserves the exact environment and commands required to reproduce the demo.
- Save README.md.
You should see sections for project behavior and macOS setup in the editor.
Markdown formatting looks uneven?
Check that each heading begins on its own line. Confirm that every opening code fence has a matching closing code fence.
Ask for a formatting review if the sections run together: help me inspect the opening Markdown in README.md
- Add the recording credits and local data flow below the existing setup section:
## Reference catalog and licenses
- `brahms`: “Hungarian Dance number 5,” Johannes Brahms, performed by the US Army Strings, marked CC-PDM-1.0.
- `vibeace`: “Vibe Ace,” Kevin MacLeod, CC-BY-4.0.
- `trumpet`: “solo trumpet 06,” Mihai Sorohan, CC-BY-4.0.
The Secret Mission uses `choice`, “Choice” by Admiral Bob featuring Snowflake, under CC-BY-NC-4.0 for this noncommercial learning exercise.
## Local data flow
1. Load and resample each reference recording.
2. Compute a spectrogram and retain local high-energy peaks.
3. Pair nearby peaks into hashes of `(anchor_frequency, target_frequency, time_delta)`.
4. Store each hash in `hash -> [(song_id, reference_time)]`.
5. Fingerprint a query and retrieve only matching posting lists.
6. Vote on `(song_id, reference_time - query_time)`.
7. Return the strongest aligned cluster with its offset, votes, and confidence.
Why include credits and data flow?
The credits preserve the documented title, attribution, and license for each recording. The numbered flow maps each local function to its role in recognition.
- Save README.md.
You should now see four catalog keys and a seven-stage local recognition flow.
Data flow list does not render in order?
Place each numbered item on its own line. Keep a blank line between the license paragraph and the local data flow heading.
Use this prompt if the Markdown list still breaks: help me fix the numbered data flow in README.md
- Add the production architecture below the local data flow:
## Production-style architecture
```mermaid
flowchart LR
Catalog[Licensed catalog] --> Ingest[Offline ingestion workers]
Ingest --> FP[Fingerprint workers]
FP --> Router[Hash partition router]
Router --> S1[Index shard 1]
Router --> S2[Index shard 2]
Router --> SN[Index shard N]
Client[Mobile or web client] --> API[API ingress]
API --> QFP[Online fingerprint workers]
QFP --> QRouter[Parallel hash lookups]
QRouter --> S1
QRouter --> S2
QRouter --> SN
S1 --> Votes[Vote aggregator]
S2 --> Votes
SN --> Votes
Votes --> Meta[Metadata cache or store]
Meta --> API
```
How does the architecture split work?
The offline path converts licensed catalog audio into sharded fingerprint postings. The online path fingerprints a client query and fans its hashes out to those same shards.
The vote aggregator reconstructs coherent song-and-offset evidence. The metadata store keeps descriptive records separate from the search payload.
- Save README.md.
You should see an offline catalog path and an online client path converge on the shared index shards.
Architecture block not recognized?
Confirm that the opening fence begins with ```mermaid and the final fence contains three backticks. Keep every connector inside those fences.
Ask for help with the diagram syntax if needed: help me inspect the Mermaid architecture block
- Add the index, partitioning, and storage decisions below the architecture diagram:
## Scaling decisions
### Inverted indexing
The query does not compare itself with every full song. Each fingerprint hash addresses a posting list containing only references that share that hash. This shifts work from catalog scans to sparse lookups plus vote aggregation.
### Partitioning and fan-out
Use the hash as the partition key so catalog ingestion and online queries route the same key to the same shard. A query contains many hashes, so its lookups can fan out in parallel. The vote aggregator combines the returned postings by song and time offset.
### Storage efficiency
Store compact hash keys and postings instead of raw audio in the online search tier. Keep metadata in a separate store keyed by song ID. The local benchmark reports raw sample count, unique key count, and posting count without pretending Python object counts are production byte estimates.
Why shard by fingerprint hash?
The same hash reaches the same shard during catalog ingestion and query lookup. This makes the fingerprint key the routing contract for both paths.
Parallel fan-out lowers retrieval time while the vote aggregator combines evidence returned by multiple shards.
- Save README.md.
You should now see separate scaling decisions for indexing, partitioning, and storage.
Scaling sections blend together?
Keep each third-level heading on its own line. Leave one blank line before each explanation paragraph.
Use this prompt to inspect the hierarchy: help me fix the scaling headings in README.md
- Add the latency, noise, and evaluation sections below the storage discussion:
### Latency
Measure fingerprint extraction separately from index lookup. Fingerprinting is stateless CPU work that can scale horizontally. Lookup latency depends on posting-list fan-out, shard balance, network round trips, and the cost of aggregating votes. Production monitoring should report latency percentiles rather than only a mean.
### Noise tolerance and false positives
Prominent spectral peaks may survive when noise hides weaker spectrogram cells. A real match creates many hashes that agree on one time offset. Random collisions tend to scatter across songs and offsets. Confidence thresholds trade recall for fewer false positives and must be calibrated on both known and unknown recordings.
## Reproducible evaluation
`evaluate.py` uses fixed seeds, three relative offsets, and SNR values of 30, 15, and 5 dB for every reference track. It reports:
- Known-query accuracy for both matchers
- False-positive rate on deterministic noise probes
- Baseline comparison latency
- Fingerprint extraction time
- Inverted-index lookup latency
- Raw sample, unique hash, and posting counts
`test_recognizer.py` verifies a shifted noisy `vibeace` query and confirms that the index contains postings.
What production signals matter?
Latency percentiles reveal slow requests hidden by a mean. Threshold-dependent false positives reveal the cost of promoting random collisions.
The reproducible evaluation section connects those concerns to the exact local measurements and tests.
- Save README.md.
You should see the benchmark's accuracy, false-positive, timing, and index-shape outputs listed under reproducible evaluation.
Evaluation list renders as a paragraph?
Leave a blank line after the sentence ending in It reports:. Keep each measurement on a separate line beginning with a hyphen.
Ask for a list-format review if needed: help me fix the evaluation list in README.md
- Complete the README by adding the limitations section at the bottom:
## Limitations
- The catalog is intentionally tiny and in memory.
- Hashes are Python tuples rather than packed integers.
- Thresholds are educational and must be recalibrated for a larger corpus.
- The matcher is not designed for pitch shifts, tempo changes, live covers, or humming.
- The benchmark measures one Mac process, not distributed network latency.
Why state the limitations?
The benchmark proves behavior inside a tiny in-memory catalog on one Mac. The limitations prevent those local results from being mistaken for distributed production guarantees.
- Save README.md.
You should see five explicit limitations at the bottom of the document.
README ends before the limitations?
Scroll to the final line of README.md. Confirm that all five limitation bullets appear below the heading.
Use this prompt if part of the file was overwritten: help me compare my README.md section order
✔️ Awesome, I've got everything!
Your README now connects the local recognizer to a sharded production architecture and documents every reproducible command.
ⓧ I'd like to double check the full code
Compare your saved README.md with this complete version.
# Mini-Shazam
A small, explainable recording-identification system built to connect audio fingerprints with production system-design concepts.
## What it demonstrates
- A direct waveform baseline fails when a recording starts at a different offset.
- Sparse spectrogram peaks preserve useful evidence under added noise.
- Peak pairs create more discriminative hashes than individual peaks.
- An inverted index maps each hash directly to candidate song and time postings.
- Time-offset voting distinguishes a coherent match from random collisions.
- Deterministic evaluation separates feature-extraction cost from index lookup latency.
This is recording identification, not melody recognition. It expects a degraded excerpt of the same underlying recording and is not designed to identify humming, cover versions, or changed pitch and tempo.
## Setup on macOS
Python 3.12 or newer is required by the pinned scientific packages.
```bash
python -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt
```
Run the demo, benchmark, and tests:
```bash
python recognizer.py
python evaluate.py
python -m unittest -v
```
The first run downloads and caches the documented librosa example recordings. The demo creates `constellation.png`.
## Reference catalog and licenses
- `brahms`: “Hungarian Dance number 5,” Johannes Brahms, performed by the US Army Strings, marked CC-PDM-1.0.
- `vibeace`: “Vibe Ace,” Kevin MacLeod, CC-BY-4.0.
- `trumpet`: “solo trumpet 06,” Mihai Sorohan, CC-BY-4.0.
The Secret Mission uses `choice`, “Choice” by Admiral Bob featuring Snowflake, under CC-BY-NC-4.0 for this noncommercial learning exercise.
## Local data flow
1. Load and resample each reference recording.
2. Compute a spectrogram and retain local high-energy peaks.
3. Pair nearby peaks into hashes of `(anchor_frequency, target_frequency, time_delta)`.
4. Store each hash in `hash -> [(song_id, reference_time)]`.
5. Fingerprint a query and retrieve only matching posting lists.
6. Vote on `(song_id, reference_time - query_time)`.
7. Return the strongest aligned cluster with its offset, votes, and confidence.
## Production-style architecture
```mermaid
flowchart LR
Catalog[Licensed catalog] --> Ingest[Offline ingestion workers]
Ingest --> FP[Fingerprint workers]
FP --> Router[Hash partition router]
Router --> S1[Index shard 1]
Router --> S2[Index shard 2]
Router --> SN[Index shard N]
Client[Mobile or web client] --> API[API ingress]
API --> QFP[Online fingerprint workers]
QFP --> QRouter[Parallel hash lookups]
QRouter --> S1
QRouter --> S2
QRouter --> SN
S1 --> Votes[Vote aggregator]
S2 --> Votes
SN --> Votes
Votes --> Meta[Metadata cache or store]
Meta --> API
```
## Scaling decisions
### Inverted indexing
The query does not compare itself with every full song. Each fingerprint hash addresses a posting list containing only references that share that hash. This shifts work from catalog scans to sparse lookups plus vote aggregation.
### Partitioning and fan-out
Use the hash as the partition key so catalog ingestion and online queries route the same key to the same shard. A query contains many hashes, so its lookups can fan out in parallel. The vote aggregator combines the returned postings by song and time offset.
### Storage efficiency
Store compact hash keys and postings instead of raw audio in the online search tier. Keep metadata in a separate store keyed by song ID. The local benchmark reports raw sample count, unique key count, and posting count without pretending Python object counts are production byte estimates.
### Latency
Measure fingerprint extraction separately from index lookup. Fingerprinting is stateless CPU work that can scale horizontally. Lookup latency depends on posting-list fan-out, shard balance, network round trips, and the cost of aggregating votes. Production monitoring should report latency percentiles rather than only a mean.
### Noise tolerance and false positives
Prominent spectral peaks may survive when noise hides weaker spectrogram cells. A real match creates many hashes that agree on one time offset. Random collisions tend to scatter across songs and offsets. Confidence thresholds trade recall for fewer false positives and must be calibrated on both known and unknown recordings.
## Reproducible evaluation
`evaluate.py` uses fixed seeds, three relative offsets, and SNR values of 30, 15, and 5 dB for every reference track. It reports:
- Known-query accuracy for both matchers
- False-positive rate on deterministic noise probes
- Baseline comparison latency
- Fingerprint extraction time
- Inverted-index lookup latency
- Raw sample, unique hash, and posting counts
`test_recognizer.py` verifies a shifted noisy `vibeace` query and confirms that the index contains postings.
## Limitations
- The catalog is intentionally tiny and in memory.
- Hashes are Python tuples rather than packed integers.
- Thresholds are educational and must be recalibrated for a larger corpus.
- The matcher is not designed for pitch shifts, tempo changes, live covers, or humming.
- The benchmark measures one Mac process, not distributed network latency.
Before the final check, predict whether the benchmark and test suite will agree on the known noisy vibeace query. Keep your answer in mind.
- Run the final benchmark and core tests with these commands:
python evaluate.py
python -m unittest -v
What does the final check prove?
The benchmark should print both method summaries, false-positive rates, timing measurements, and index-shape counts. The tests should report the known noisy clip as correctly identified.
A final OK confirms that the core recognition and index guarantees still hold after the evaluation work.
Final benchmark or tests failing?
Run each command separately to identify which file needs attention. Compare that file with its full-code tab before changing recognizer.py.
Share the failing command and traceback through this prompt: help me diagnose the final Mini-Shazam benchmark or test failure
That completes the evidence behind your recognizer. You now have repeatable measurements, passing regression tests, and a production design that explains where the system can scale.
Secret mission
Reject an Unknown Track
A raw fingerprint winner can come from accidental hash collisions. In this mission, you will add aligned-vote and confidence thresholds so the recognizer accepts the known clip while rejecting the out-of-catalog `choice` recording.
Clean Up Your Resources
Clean Up Your Resources
Everything in this project runs locally, so there are no ongoing cloud or API costs. Decide whether to keep your resources available, pause your local workspace, or delete the project entirely.
Resources you used:
- Your local project folder is open in Visual Studio Code.
- Core recognition files include recognizer.py and evaluate.py.
- Test files include test_recognizer.py and test_thresholds.py.
- Secret Mission files include threshold_recognizer.py and evaluate_unknown.py.
- Supporting resources include .venv, requirements.txt, README.md, and constellation.png.
Keep everything running
No action needed. Choose this if you want to keep benchmarking the recognizer or calibrating its confidence thresholds.
- Keep the existing project folder on your Mac.
- Leave the existing .venv interpreter selected in Visual Studio Code.
- Retain constellation.png as the visual result of your peak-extraction pipeline.
Pause - I'll come back to this later
Pausing ends the active local shell while keeping every project file intact. You can resume from the same folder later.
- Save any open files in Visual Studio Code.
- Close the integrated terminal that shows .venv as active.
- Close Visual Studio Code.
- Return to the same project folder in Visual Studio Code when you are ready to continue.
- Select the existing .venv interpreter again.
Delete - I don't want to use this again
Deletion removes the project code and its isolated environment from your Mac. Permanently clearing the deleted folder makes this cleanup irreversible.
- Copy any project files you want to keep into another folder.
- Find the top-level folder name displayed above recognizer.py in Visual Studio Code.
- Record that folder name here: your mini-Shazam project folder.
- Close the integrated terminal that shows .venv as active.
- Close Visual Studio Code.
Your workspace is now closed. The macOS file browser can remove its folder cleanly.
- Use the macOS file browser to search for your mini-Shazam project folder.
- Confirm the selected folder contains recognizer.py.
- Delete your mini-Shazam project folder from its current location.
- Permanently clear the deleted folder from your Mac.
You now have a clean local slate. The project workspace and its isolated Python environment are removed from your Mac.
Nice Work!
Nice Work!
You made it. Your Mini-Shazam now identifies shifted noisy recordings through explainable audio fingerprints while exposing the system-design choices behind each result.
You've learned how to:
- Expose the alignment limits of waveform template matching. Transform noisy audio into a sparse spectrogram. Retain local spectral peaks that survive shifted starts plus added noise.
- Create discriminative audio fingerprints from bounded peak pairs. Retrieve matching postings through an inverted index. Reconstruct the recording plus its alignment through time-offset voting.
- Benchmark recognition accuracy across controlled offsets plus noise levels. Measure false positives. Separate feature-extraction time from lookup latency. Map the local recognizer to a sharded production architecture.
- Secret Mission: Added aligned-vote thresholds plus confidence thresholds. The recognizer now rejects the out-of-catalog choice clip as No confident match.
Ready to quiz yourself?