Build a Personalized Sign Recognizer
Train a webcam app to recognize your own handshapes in real time.
Introduction
30 Second Summary
A webcam sees your hand as pixels. It needs examples before it can recognize the poses that matter to you.
In this project, you will build a local webcam app that uses machine learning to recognize three personalized handshapes. You will expose why hand position affects raw landmark data before making the predictions more stable.
What You'll Build
Hold a pose in front of your webcam and watch your recognizer display its predicted label with the proportion of nearby examples that voted for it.
By the end of this project, you'll have:
- A personalized training dataset containing 20 recorded examples for each of three handshapes.
- A live handshape recognizer that displays one of your learned labels with its neighbor vote.
- A feature-engineering comparison that shows how normalization makes predictions more stable as your hand moves or changes size.
- Secret Mission: Mark every prediction without a unanimous neighbor vote as uncertain.
Are there any prerequisites?
You need a Windows computer with a working webcam. Basic Python syntax is enough because the project guides you through the package setup.
Before We Start
Before any hands-on work begins, this step locks in the purpose of your personalized recognizer. Your plan will connect three distinct static poses to the neutral labels sign_1, sign_2, and sign_3.
Set Up the Windows ML Workspace
Your personalized recognizer relies on several native Python packages. A clean environment prevents their dependencies from colliding.
You'll use Visual Studio Code to create a virtual environment for the project. Pinned dependencies then make the setup reproducible.
In this step, get ready to:
- Verify that your Python version supports the pinned packages.
- Create an isolated environment with the exact project dependencies.
- Download the official Hand Landmarker model.
Verify Python and isolate the project
The pinned scikit-learn release requires Python 3.11 or newer. Confirming the interpreter first prevents an incompatible package installation.
- Press the Windows key to open Windows search.
- Type Visual Studio Code into the search field.
- Press Enter to open Visual Studio Code.
- Click File in the menu bar.
- Click Open Folder.
- Select the project folder for your personalized recognizer.
- Click Select Folder.
You'll see the empty project folder in the Visual Studio Code sidebar. This is where every project file belongs.
- Click View in the menu bar.
- Click Terminal to open the integrated terminal.
- Check the available Python version by running this command:
python --version
What does this command check?
The command asks the active Python executable to report its version. The result determines whether the pinned packages can run in this environment.
✔️ I see a supported Python version
Your result is between Python 3.11 and 3.14. Your interpreter is compatible with the project packages.
ⓧ I see an older Python version
Your current interpreter is below the project's compatibility floor. Install Python 3.14.8 before creating the environment.
- Open the official Python 3.14.8 Windows release page.
- Download the Windows installer for your computer.
- Run the downloaded installer.
- Approve the Windows permission prompt if it appears.
- Complete the installation using the default settings.
- Return to Visual Studio Code after the installer finishes.
- Reopen the integrated terminal from the View menu.
- Check the new Python version by running this command:
python --version
What does this command confirm?
The second check confirms that the reopened terminal can find the newly installed interpreter. It should now report Python 3.14.8.
ⓧ The command is not found
Windows cannot currently find a Python executable. Installing Python 3.14.8 provides the interpreter required for this project.
- Open the official Python 3.14.8 Windows release page.
- Download the Windows installer for your computer.
- Run the downloaded installer.
- Approve the Windows permission prompt if it appears.
- Complete the installation using the default settings.
- Return to Visual Studio Code after the installer finishes.
- Reopen the integrated terminal from the View menu.
- Confirm that Windows can now find Python by running this command:
python --version
What does this command confirm?
The reopened terminal should now report Python 3.14.8. That output confirms the installation is available to your project.
A virtual environment keeps this project's packages separate from other Python projects on your computer. The .venv folder stores that isolated environment inside the project folder.
- Create the isolated environment by running this command:
py -m venv .venv
What does this command create?
The Python launcher creates a local environment named .venv. Packages installed after activation stay isolated inside that folder.
- Activate the new environment by running this command:
.venv\Scripts\activate
What does activation change?
Activation points the terminal at the interpreter inside .venv. The packages you install next stay attached to this project.
You'll see (.venv) at the beginning of the terminal prompt. That prefix confirms the isolated environment is active.
Don't see the environment prefix?
- Confirm that the terminal is inside the project folder shown in the Visual Studio Code sidebar.
- Check that the .venv folder appears in the sidebar.
- Run the activation command again after the environment folder finishes appearing.
Still stuck? Help me activate my Windows Python virtual environment
Install the pinned packages
The dependency file gives every installation the same versions of MediaPipe, OpenCV, plus scikit-learn. Exact pins remove version drift from the rest of the project.
- Select the new-file control in the Visual Studio Code Explorer sidebar.
- Name the file requirements.txt.
- Add the exact dependency pins by pasting this content into requirements.txt:
mediapipe==1.1.0
opencv-contrib-python==5.0.0.93
scikit-learn==1.9.1
What do these pins control?
- MediaPipe provides the pretrained hand landmark detector used by the webcam pipeline.
- The OpenCV contrib wheel provides camera access through the shared cv2 namespace.
- Scikit-learn provides the classifier used later in the project.
- Save requirements.txt.
- Install the pinned packages into the active environment by running:
python -m pip install -r requirements.txt
What does this command install?
Python reads each exact pin from requirements.txt. The active environment receives those versions without changing other Python projects.
The first installation may take a few minutes while Windows downloads the package wheels. You'll know it is complete when the terminal returns to the (.venv) prompt without an installation error.
Package installation failed?
- Confirm that the terminal prompt begins with (.venv).
- Compare every package name in requirements.txt with the full-file reference below.
- Keep only the pinned OpenCV contrib wheel in this environment because OpenCV wheel variants share the cv2 namespace.
Still stuck? Help me diagnose my pinned Python package installation
✔️ Awesome, I've got everything!
Your dependency file is saved. The three pinned packages are installed inside .venv.
ⓧ I'd like to double check the full code
mediapipe==1.1.0
opencv-contrib-python==5.0.0.93
scikit-learn==1.9.1
Download the Hand Landmarker model
The Hand Landmarker model converts a camera frame into 21 landmark positions. A small download script stores the official model at the path expected by the later webcam code.
- Select the new-file control in the Visual Studio Code Explorer sidebar.
- Name the file download_model.py.
- Add the model download logic by pasting this code into download_model.py:
from pathlib import Path
from urllib.request import urlretrieve
MODEL_URL = (
"https://storage.googleapis.com/mediapipe-models/hand_landmarker/"
"hand_landmarker/float16/1/hand_landmarker.task"
)
MODEL_PATH = Path("models") / "hand_landmarker.task"
def main():
MODEL_PATH.parent.mkdir(exist_ok=True)
if MODEL_PATH.exists():
print(f"Model already exists at {MODEL_PATH}")
return
print("Downloading the MediaPipe Hand Landmarker model...")
urlretrieve(MODEL_URL, MODEL_PATH)
print(f"Downloaded model to {MODEL_PATH}")
if __name__ == "__main__":
main()
What does this code do?
- The MODEL_URL constant points to the official model file.
- The MODEL_PATH constant keeps the model inside the project's models folder.
- The main() function creates the destination folder when needed.
- The existence check prevents an unnecessary second download.
- Save download_model.py.
✔️ Awesome, I've got everything!
Your download script is saved. It is ready to place the model in the expected project folder.
ⓧ I'd like to double check the full code
from pathlib import Path
from urllib.request import urlretrieve
MODEL_URL = (
"https://storage.googleapis.com/mediapipe-models/hand_landmarker/"
"hand_landmarker/float16/1/hand_landmarker.task"
)
MODEL_PATH = Path("models") / "hand_landmarker.task"
def main():
MODEL_PATH.parent.mkdir(exist_ok=True)
if MODEL_PATH.exists():
print(f"Model already exists at {MODEL_PATH}")
return
print("Downloading the MediaPipe Hand Landmarker model...")
urlretrieve(MODEL_URL, MODEL_PATH)
print(f"Downloaded model to {MODEL_PATH}")
if __name__ == "__main__":
main()
Before you run the script, what message do you expect after the model reaches the models folder?
- Download the official model by running this command:
python download_model.py
What does this command run?
Python executes the saved download script inside the active environment. The script creates the models folder before saving the model file.
You'll see Downloaded model to models\hand_landmarker.task. A later run may show Model already exists at models\hand_landmarker.task instead.
Model download did not finish?
- Confirm that download_model.py is saved inside the open project folder.
- Check that the terminal still shows the (.venv) prefix.
- Confirm that your computer has an active internet connection before rerunning the command.
Still stuck? Help me troubleshoot the Hand Landmarker model download
That's the foundation in place. Your isolated workspace now has the pinned packages plus the official model needed for real-time landmark detection.
Next, you'll connect this workspace to your webcam. You'll see your hand represented by 21 live landmark points.
See Your Hand as Landmarks
Your Python environment now has the packages and model it needs. Before you collect training data, you need visible proof that the webcam can consistently find your hand.
In this step, you will connect MediaPipe to OpenCV. The detected landmarks will become a 63-value feature vector for each hand.
In this step, get ready to:
- Build a reusable helper that detects one hand.
- Create a webcam preview with 21 green landmark points.
- Verify that the points follow your hand as it moves.
Build the reusable landmark helper
The Hand Landmarker returns 21 points with x, y, and z coordinates. Limiting detection to one hand gives every sample the same 63-value shape.
Why use landmarks instead of images?
Landmarks reduce each camera frame to compact hand geometry. This keeps the project focused on supervised learning and feature engineering.
A custom convolutional neural network would require image labeling plus a longer training process. Landmark features let your small personalized dataset train quickly.
- In the file sidebar in Visual Studio Code, click the new-file button.
- Type hand_features.py as the file name.
- Press Enter to create the file.
- Add the path, timing, camera, and MediaPipe Tasks imports by pasting this at the top of hand_features.py:
from pathlib import Path
import time
import cv2 as cv
from mediapipe.tasks import python
from mediapipe.tasks.python import vision
from mediapipe.tasks.python.vision.core import image as image_lib
What do these imports provide?
- Path builds a platform-safe location for the downloaded model.
- The time module provides increasing timestamps for video detection.
- The cv module handles camera frames and drawing.
- The MediaPipe Tasks imports configure the Hand Landmarker and its input image.
- Save hand_features.py.
- Check the imports by running this command in the terminal from earlier:
python -m py_compile hand_features.py
What confirms success?
This command compiles the file without running the webcam. A successful check returns the terminal prompt without a traceback.
Are the imports failing to compile?
Check each import against the snippet above. Confirm that the terminal still shows your activated .venv environment.
Still stuck? Help me troubleshoot the imports in hand_features.py.
- Define the model location and Hand Landmarker builder by adding this below the imports:
MODEL_PATH = Path("models") / "hand_landmarker.task"
_last_timestamp_ms = 0
def create_landmarker():
if not MODEL_PATH.exists():
raise FileNotFoundError(
"Missing models/hand_landmarker.task. Run python download_model.py first."
)
options = vision.HandLandmarkerOptions(
base_options=python.BaseOptions(model_asset_path=str(MODEL_PATH)),
running_mode=vision.RunningMode.VIDEO,
num_hands=1,
)
return vision.HandLandmarker.create_from_options(options)
What does this code do?
- MODEL_PATH points to the task model you downloaded in the previous step.
- create_landmarker checks that the model exists before configuring detection.
- VIDEO mode prepares the detector for a sequence of webcam frames.
- The one-hand limit keeps every future feature row at 63 values.
- Save hand_features.py.
- Check the new configuration by running:
python -m py_compile hand_features.py
What confirms success?
The compiler checks the new constants plus the complete function. You will return to the terminal prompt when their syntax is valid.
Is the configuration failing to compile?
Check the indentation inside create_landmarker(). Confirm that each opening parenthesis has a matching closing parenthesis.
Still stuck? Help me fix the create_landmarker function.
- Add the video timestamp helper below create_landmarker() by pasting:
def next_timestamp_ms():
global _last_timestamp_ms
current_timestamp_ms = time.monotonic_ns() // 1_000_000
_last_timestamp_ms = max(current_timestamp_ms, _last_timestamp_ms + 1)
return _last_timestamp_ms
Why does video detection need this?
MediaPipe VIDEO mode expects each frame to have a millisecond timestamp that is greater than the previous one. This helper protects that ordering even when two frames are processed within the same millisecond.
- Save hand_features.py.
- Check the timestamp helper by running:
python -m py_compile hand_features.py
What confirms success?
The compiler verifies the global assignment plus the timestamp calculation. The terminal prompt returns when the function is valid.
Is the timestamp helper failing?
Confirm that _last_timestamp_ms has the same spelling in the constant and function. Keep the leading underscore in both places.
Still stuck? Help me debug next_timestamp_ms.
- Convert each camera frame into a MediaPipe image by adding this below next_timestamp_ms():
def detect_hand(landmarker, frame):
rgb_frame = cv.cvtColor(frame, cv.COLOR_BGR2RGB)
image = image_lib.Image(
image_format=image_lib.ImageFormat.SRGB,
data=rgb_frame,
)
result = landmarker.detect_for_video(image, next_timestamp_ms())
if not result.hand_landmarks:
return None
return result.hand_landmarks[0]
What happens to each frame?
- OpenCV supplies a BGR camera frame.
- The color conversion produces the RGB data expected by the MediaPipe image.
- detect_for_video processes the frame with the next timestamp.
- The function returns one detected hand or None when no hand is visible.
- Save hand_features.py.
- Check the frame conversion and detection logic by running:
python -m py_compile hand_features.py
What confirms success?
The compiler checks the image constructor plus the detection branches. A valid file returns you to the terminal prompt.
Is hand detection failing to compile?
Check that image_lib.Image uses the imported image_lib name. Confirm that the return statements align with their matching conditions.
Still stuck? Help me fix detect_hand.
- Add the landmark drawing helper below detect_hand() by pasting:
def draw_landmarks(frame, landmarks):
height, width = frame.shape[:2]
for landmark in landmarks:
x = min(max(int(landmark.x * width), 0), width - 1)
y = min(max(int(landmark.y * height), 0), height - 1)
cv.circle(frame, (x, y), 4, (0, 255, 0), -1)
How are the points drawn?
The landmark coordinates use frame-relative values. Multiplying by the frame dimensions converts them into pixel positions.
The bounds keep every point inside the image. Each point becomes a filled green circle with a four-pixel radius.
- Save hand_features.py.
- Check the drawing helper by running:
python -m py_compile hand_features.py
What confirms success?
The compiler checks the coordinate calculations plus the drawing call. The terminal prompt returns when their syntax is valid.
Is the drawing helper failing?
Confirm that height, width matches the order shown above. Check the parentheses around both bounded coordinate calculations.
Still stuck? Help me debug draw_landmarks.
- Turn the detected points into model-ready numbers by adding this below draw_landmarks():
def raw_features(landmarks):
features = []
for landmark in landmarks:
features.extend(
(float(landmark.x), float(landmark.y), float(landmark.z))
)
return features
How does this become a feature vector?
The loop preserves the detector's landmark order. Each point contributes three floating-point coordinates.
Twenty-one points produce 63 values in one flat list. Every future training row will use this shape.
- Save hand_features.py.
- Check the raw feature helper by running:
python -m py_compile hand_features.py
What confirms success?
The compiler checks the loop plus the three-coordinate tuple. You will return to the terminal prompt when the helper is valid.
Is raw feature extraction failing?
Confirm that features.extend() contains all three landmark coordinates. Keep return features outside the loop.
Still stuck? Help me fix raw_features.
- Prepare the normalization function location by adding this final check below raw_features():
def normalize_features(features):
if len(features) != 63:
raise ValueError("Expected 63 values from 21 three-dimensional landmarks.")
Why prepare this function now?
The function reserves the feature-normalization entry point used later in the project. Its current length check protects the fixed 63-value shape.
- Save hand_features.py.
- Check the complete helper file by running:
python -m py_compile hand_features.py
What confirms success?
The compiler now checks every helper in the file. A successful result returns the terminal prompt without a traceback.
Is the complete helper failing?
Use the full-file reference below to compare each function in order. Pay close attention to indentation between neighboring functions.
Still stuck? Help me compare my complete hand_features.py file.
✔️ Awesome, I've got everything!
Your helper file now loads the Hand Landmarker and turns one detected hand into a fixed-size feature vector.
ⓧ I'd like to double check the full code
from pathlib import Path
import time
import cv2 as cv
from mediapipe.tasks import python
from mediapipe.tasks.python import vision
from mediapipe.tasks.python.vision.core import image as image_lib
MODEL_PATH = Path("models") / "hand_landmarker.task"
_last_timestamp_ms = 0
def create_landmarker():
if not MODEL_PATH.exists():
raise FileNotFoundError(
"Missing models/hand_landmarker.task. Run python download_model.py first."
)
options = vision.HandLandmarkerOptions(
base_options=python.BaseOptions(model_asset_path=str(MODEL_PATH)),
running_mode=vision.RunningMode.VIDEO,
num_hands=1,
)
return vision.HandLandmarker.create_from_options(options)
def next_timestamp_ms():
global _last_timestamp_ms
current_timestamp_ms = time.monotonic_ns() // 1_000_000
_last_timestamp_ms = max(current_timestamp_ms, _last_timestamp_ms + 1)
return _last_timestamp_ms
def detect_hand(landmarker, frame):
rgb_frame = cv.cvtColor(frame, cv.COLOR_BGR2RGB)
image = image_lib.Image(
image_format=image_lib.ImageFormat.SRGB,
data=rgb_frame,
)
result = landmarker.detect_for_video(image, next_timestamp_ms())
if not result.hand_landmarks:
return None
return result.hand_landmarks[0]
def draw_landmarks(frame, landmarks):
height, width = frame.shape[:2]
for landmark in landmarks:
x = min(max(int(landmark.x * width), 0), width - 1)
y = min(max(int(landmark.y * height), 0), height - 1)
cv.circle(frame, (x, y), 4, (0, 255, 0), -1)
def raw_features(landmarks):
features = []
for landmark in landmarks:
features.extend(
(float(landmark.x), float(landmark.y), float(landmark.z))
)
return features
def normalize_features(features):
if len(features) != 63:
raise ValueError("Expected 63 values from 21 three-dimensional landmarks.")
Create and run the visible preview
The helper can now process a frame. The preview supplies those frames continuously while keeping webcam cleanup inside a finally block.
- In the file sidebar, click the new-file button.
- Type preview_landmarks.py as the file name.
- Press Enter to create the file.
- Build the webcam loop by pasting this into preview_landmarks.py:
import cv2 as cv
from hand_features import create_landmarker, detect_hand, draw_landmarks
def main():
capture = cv.VideoCapture(0)
if not capture.isOpened():
raise RuntimeError("Could not open webcam 0.")
try:
with create_landmarker() as landmarker:
while True:
frame_ok, frame = capture.read()
if not frame_ok:
raise RuntimeError("Could not read a frame from the webcam.")
cv.imshow("Hand Landmark Preview", frame)
if cv.waitKey(1) == ord("q"):
break
finally:
capture.release()
cv.destroyAllWindows()
if __name__ == "__main__":
main()
What does this first preview do?
- VideoCapture opens camera index zero.
- The loop reads one frame at a time and displays it in a preview window.
- The q key exits the loop.
- The cleanup block releases the camera and closes every OpenCV window.
The command occupies the terminal while the preview window is open. Pressing q inside that window returns control to the terminal.
- Save preview_landmarks.py.
- Start the basic webcam loop by running:
python preview_landmarks.py
What should you see?
You should see your live camera feed in a window named Hand Landmark Preview. This first run proves that camera zero opens successfully.
- Click the preview window to give it keyboard focus.
- Press q to close the preview.
Does the webcam window stay closed?
Close any other application that is using the webcam. Confirm that Windows allows Python to access the camera.
Still stuck? Help me open webcam zero with OpenCV.
- Return to preview_landmarks.py.
- Find the line containing cv.imshow("Hand Landmark Preview", frame).
- Insert the detection and overlay logic directly above that line by pasting:
landmarks = detect_hand(landmarker, frame)
if landmarks:
draw_landmarks(frame, landmarks)
status = f"{len(landmarks)} landmarks detected"
else:
status = "Show one hand"
cv.putText(
frame,
status,
(20, 35),
cv.FONT_HERSHEY_SIMPLEX,
0.8,
(0, 255, 0),
2,
cv.LINE_AA,
)
How does the overlay work?
- detect_hand sends the current frame through the Hand Landmarker.
- draw_landmarks adds one green point for every detected landmark.
- The status line reports the number of detected points.
- A frame without a detected hand displays an instruction to show one hand.
- Save preview_landmarks.py.
✔️ Awesome, I've got everything!
Your preview now connects the webcam loop to the reusable landmark helpers.
ⓧ I'd like to double check the full code
import cv2 as cv
from hand_features import create_landmarker, detect_hand, draw_landmarks
def main():
capture = cv.VideoCapture(0)
if not capture.isOpened():
raise RuntimeError("Could not open webcam 0.")
try:
with create_landmarker() as landmarker:
while True:
frame_ok, frame = capture.read()
if not frame_ok:
raise RuntimeError("Could not read a frame from the webcam.")
landmarks = detect_hand(landmarker, frame)
if landmarks:
draw_landmarks(frame, landmarks)
status = f"{len(landmarks)} landmarks detected"
else:
status = "Show one hand"
cv.putText(
frame,
status,
(20, 35),
cv.FONT_HERSHEY_SIMPLEX,
0.8,
(0, 255, 0),
2,
cv.LINE_AA,
)
cv.imshow("Hand Landmark Preview", frame)
if cv.waitKey(1) == ord("q"):
break
finally:
capture.release()
cv.destroyAllWindows()
if __name__ == "__main__":
main()
Before you run this, make a quick prediction about whether all 21 points will stay attached as you move your hand.
- Start the completed landmark preview by running:
python preview_landmarks.py
What should you see?
Show one hand to the camera. You should see 21 green points plus the status 21 landmarks detected.
Move your hand around the frame. The points should follow your fingers and palm in real time.
Do the landmarks disappear or freeze?
Move into brighter lighting and keep your full hand inside the frame. Hold one hand toward the camera with your fingers clearly separated.
Still stuck? Help me troubleshoot the Hand Landmark Preview.
- Click the preview window to give it keyboard focus.
- Press q to close the preview cleanly.
That is the camera pipeline working: your preview now detects one hand and follows it with 21 landmark points. The terminal prompt returns after the webcam is released.
Next, you will capture those 63-value feature vectors and pair them with three personalized labels.
Collect Three Labeled Handshape Datasets
The landmark preview proved that MediaPipe can locate your hand. A supervised classifier still needs examples paired with the labels you want it to predict.
Those pairs form labeled data. In this step, you will build a keyboard-controlled collector before recording three balanced classes.
In this step, get ready to:
- Build a webcam collector that saves labeled hand landmarks.
- Record 20 centered samples for each of three handshapes.
- Confirm the dataset contains 60 labeled rows.
Build the keyboard-driven collector
Each detected hand becomes a 63-value feature vector containing the 21 landmarks' coordinates. The collector stores that vector beside the label supplied at launch.
Why use neutral labels?
The labels sign_1, sign_2, and sign_3 describe personalized handshapes without assigning linguistic meaning. This keeps the project focused on isolated pose classification.
You will assemble the first working collector in two adjacent parts. The first part defines its inputs and dataset location.
- Create collect_data.py inside the project folder using the Visual Studio Code file list.
- Add the collector's imports and argument parser by copying this first block into collect_data.py:
import argparse
import csv
from pathlib import Path
import cv2 as cv
from hand_features import (
create_landmarker,
detect_hand,
draw_landmarks,
raw_features,
)
DATA_PATH = Path("data") / "landmarks.csv"
FEATURE_COUNT = 63
def parse_args():
parser = argparse.ArgumentParser(
description="Collect labeled hand landmark samples."
)
parser.add_argument("label", help="Label to save, such as sign_1")
parser.add_argument("--samples", type=int, default=20)
return parser.parse_args()
What does this code do?
- The argument parser requires a label for every collection session.
- The optional --samples argument controls the target number of saved examples.
- The DATA_PATH value sends every sample to data/landmarks.csv.
- The FEATURE_COUNT value reserves 63 columns for 21 three-dimensional landmarks.
- Add the collector's initial main() function directly below parse_args() by copying this block:
def main():
args = parse_args()
DATA_PATH.parent.mkdir(exist_ok=True)
needs_header = not DATA_PATH.exists() or DATA_PATH.stat().st_size == 0
capture = cv.VideoCapture(0)
if not capture.isOpened():
raise RuntimeError("Could not open webcam 0.")
captured = 0
try:
with DATA_PATH.open("a", newline="", encoding="utf-8") as data_file:
writer = csv.writer(data_file)
if needs_header:
writer.writerow(
["label"] + [f"feature_{index}" for index in range(FEATURE_COUNT)]
)
finally:
capture.release()
cv.destroyAllWindows()
print(f"Saved {captured} samples for {args.label} to {DATA_PATH}")
if __name__ == "__main__":
main()
How does the scaffold work?
- The script creates the data folder when it is missing.
- A new CSV file receives one label column plus feature_0 through feature_62.
- The webcam opens before collection begins.
- The finally block releases the webcam after the script finishes.
- Save collect_data.py.
- Check the dataset scaffold by running this command in the activated terminal:
python collect_data.py sign_1
What should this check prove?
This early run checks the label argument and webcam access. It also creates the dataset header before the collection loop is added.
You should see Saved 0 samples for sign_1 to data\landmarks.csv in the terminal. That output confirms the scaffold completed cleanly.
Does the scaffold stop with an error?
- Confirm the earlier preview window is closed so it no longer holds the webcam.
- Check that both code blocks are inside collect_data.py with the indentation shown.
Still stuck? Help me debug the initial collect_data.py scaffold
The remaining loop turns the scaffold into an interactive collector. Add both pieces before saving or running the file again.
- Find the closing parenthesis of the header's writer.writerow() call inside main().
- Insert this first loop block directly below that parenthesis:
with create_landmarker() as landmarker:
while captured < args.samples:
frame_ok, frame = capture.read()
if not frame_ok:
raise RuntimeError("Could not read a frame from the webcam.")
landmarks = detect_hand(landmarker, frame)
if landmarks:
draw_landmarks(frame, landmarks)
cv.putText(
frame,
f"Label: {args.label} Saved: {captured}/{args.samples}",
(20, 35),
cv.FONT_HERSHEY_SIMPLEX,
0.7,
(0, 255, 0),
2,
cv.LINE_AA,
)
What does this loop track?
- The loop continues until captured reaches the requested sample count.
- Each frame passes through the existing hand detector.
- Detected landmarks appear as green points in the webcam window.
- The status text shows the active label plus the saved sample count.
- Insert this second loop block immediately below the first cv.putText() call:
cv.putText(
frame,
"Space: save sample q: quit",
(20, 70),
cv.FONT_HERSHEY_SIMPLEX,
0.6,
(255, 255, 255),
2,
cv.LINE_AA,
)
cv.imshow("Collect Handshape Data", frame)
key = cv.waitKey(1)
if key == ord("q"):
break
if key == ord(" ") and landmarks:
writer.writerow([args.label, *raw_features(landmarks)])
data_file.flush()
captured += 1
How are samples saved?
- The webcam window displays the controls below the sample counter.
- Pressing Space saves one row only when a hand is detected.
- Each row starts with args.label before the 63 values returned by raw_features().
- Pressing q ends a session before the target is reached.
- Save the completed collect_data.py file.
✔️ Awesome, I've got everything!
Your collector now has the parser, webcam loop, keyboard controls, CSV writer, and cleanup logic. Keep collect_data.py saved before recording your poses.
ⓧ I'd like to double check the full code
import argparse
import csv
from pathlib import Path
import cv2 as cv
from hand_features import (
create_landmarker,
detect_hand,
draw_landmarks,
raw_features,
)
DATA_PATH = Path("data") / "landmarks.csv"
FEATURE_COUNT = 63
def parse_args():
parser = argparse.ArgumentParser(
description="Collect labeled hand landmark samples."
)
parser.add_argument("label", help="Label to save, such as sign_1")
parser.add_argument("--samples", type=int, default=20)
return parser.parse_args()
def main():
args = parse_args()
DATA_PATH.parent.mkdir(exist_ok=True)
needs_header = not DATA_PATH.exists() or DATA_PATH.stat().st_size == 0
capture = cv.VideoCapture(0)
if not capture.isOpened():
raise RuntimeError("Could not open webcam 0.")
captured = 0
try:
with DATA_PATH.open("a", newline="", encoding="utf-8") as data_file:
writer = csv.writer(data_file)
if needs_header:
writer.writerow(
["label"] + [f"feature_{index}" for index in range(FEATURE_COUNT)]
)
with create_landmarker() as landmarker:
while captured < args.samples:
frame_ok, frame = capture.read()
if not frame_ok:
raise RuntimeError("Could not read a frame from the webcam.")
landmarks = detect_hand(landmarker, frame)
if landmarks:
draw_landmarks(frame, landmarks)
cv.putText(
frame,
f"Label: {args.label} Saved: {captured}/{args.samples}",
(20, 35),
cv.FONT_HERSHEY_SIMPLEX,
0.7,
(0, 255, 0),
2,
cv.LINE_AA,
)
cv.putText(
frame,
"Space: save sample q: quit",
(20, 70),
cv.FONT_HERSHEY_SIMPLEX,
0.6,
(255, 255, 255),
2,
cv.LINE_AA,
)
cv.imshow("Collect Handshape Data", frame)
key = cv.waitKey(1)
if key == ord("q"):
break
if key == ord(" ") and landmarks:
writer.writerow([args.label, *raw_features(landmarks)])
data_file.flush()
captured += 1
finally:
capture.release()
cv.destroyAllWindows()
print(f"Saved {captured} samples for {args.label} to {DATA_PATH}")
if __name__ == "__main__":
main()
How to use this reference
Compare the file from top to bottom. Pay close attention to the indentation inside the nested file and landmark contexts.
Record three balanced classes
A balanced dataset gives every class the same number of examples. Keeping your hand near the center also controls the collection conditions for the raw-coordinate experiment.
- Choose three visually distinct static poses for sign_1, sign_2, and sign_3.
- Use the same hand for every collection session.
- Keep the camera at one fixed angle for every collection session.
- Keep the room lighting consistent for every collection session.
- Hold the first pose near the center of the webcam frame.
- Start the sign_1 session by running this command:
python collect_data.py sign_1
How does this session work?
The label argument applies sign_1 to every row saved during this run. The default target is 20 samples.
- Press Space once for each small natural variation of the first pose.
- Continue until the counter reaches 20/20.
The webcam window closes when the target is reached. You should see Saved 20 samples for sign_1 to data\landmarks.csv in the terminal.
Does the counter stay unchanged?
- Move the full hand into view until the 21 green points appear.
- Click the webcam window before pressing Space so it receives the keystroke.
- Hold each variation steady for a moment before saving it.
Need help? Help me troubleshoot a handshape sample counter that does not increase
- Hold the second pose near the center of the webcam frame.
- Start the sign_2 session by running this command:
python collect_data.py sign_2
What changes in this session?
The collector appends new rows to the existing file. Each new row receives sign_2 as its label.
- Press Space once for each small natural variation of the second pose.
- Continue until the counter reaches 20/20.
You should see Saved 20 samples for sign_2 to data\landmarks.csv after the webcam window closes.
Did the second session stop early?
- Run the sign_2 command again only if the terminal reported fewer than 20 saved samples.
- Avoid pressing q until the counter reaches the target.
Need help? Help me finish the sign_2 collection without duplicating the wrong label
Before you start the final session, predict the dataset's total row count after 20 more saves.
- Hold the third pose near the center of the webcam frame.
- Start the sign_3 session by running this command:
python collect_data.py sign_3
What completes the dataset?
This run appends the final 20 rows under sign_3. Completing it gives all three classes equal representation.
- Press Space once for each small natural variation of the third pose.
- Continue until the counter reaches 20/20.
You should see Saved 20 samples for sign_3 to data\landmarks.csv after the final session closes.
Is the final saved count below 20?
- Check that green landmarks remain visible before each press of Space.
- Restart the sign_3 session only if the previous run saved no rows.
Need help? Help me verify why the final collection session saved fewer samples than expected
- Return to the Visual Studio Code file list.
- Expand the data folder.
- Select landmarks.csv.
- Scroll to the final row in the file.
The final row should be on line 61 because line 1 contains the header. The three terminal messages confirm that each label contributed 20 rows.
Strong work. Your personalized dataset now contains 60 hand observations ready for training.
Your three balanced classes are ready. Next, you will train a k-nearest neighbors model and test what raw coordinates do when your hand moves across the frame.
Train the Position-Sensitive Classifier
Your 60 labeled examples now form a balanced dataset for three personalized handshapes. The next goal is to turn those rows into a saved classifier.
A k-nearest neighbors classifier compares each hand with nearby recorded examples. This step tests whether centered raw coordinates remain reliable when the same pose moves around the frame.
In this step, get ready to:
- Build a training script with a stratified holdout evaluation.
- Save a raw-coordinate classifier trained on all 60 samples.
- Test the saved model through a live webcam prediction loop.
Train and save the naive model
A holdout evaluation reserves part of your dataset for a same-distribution check. The final classifier then trains on all 60 rows before model persistence saves it for the webcam application.
- Create `train_model.py` in the Visual Studio Code Explorer sidebar.
- Add the imports, file paths, and mode argument by pasting this first section:
import argparse
from collections import Counter
import csv
from pathlib import Path
import pickle
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from hand_features import normalize_features
DATA_PATH = Path("data") / "landmarks.csv"
MODEL_PATH = Path("models") / "sign_knn.pkl"
def parse_args():
parser = argparse.ArgumentParser(description="Train the handshape classifier.")
parser.add_argument(
"--mode",
choices=("raw", "normalized"),
default="normalized",
)
return parser.parse_args()
How Do the Training Options Work?
- `DATA_PATH` points to the 60-row dataset you collected earlier.
- `MODEL_PATH` identifies where the fitted classifier payload is saved.
- `parse_args()` restricts the feature mode to `raw` or `normalized`.
- Add the dataset loader below `parse_args()` by pasting this function:
def load_dataset(mode):
if not DATA_PATH.exists():
raise FileNotFoundError("Missing data/landmarks.csv. Collect samples first.")
features = []
labels = []
with DATA_PATH.open("r", newline="", encoding="utf-8") as data_file:
reader = csv.reader(data_file)
next(reader)
for row in reader:
sample = [float(value) for value in row[1:]]
if len(sample) != 63:
raise ValueError("Every dataset row must contain 63 feature values.")
if mode == "normalized":
sample = normalize_features(sample)
labels.append(row[0])
features.append(sample)
return features, labels
How Does the Dataset Become Training Data?
- `load_dataset()` skips the CSV header before reading the recorded rows.
- Each label goes into `labels`.
- Each group of 63 coordinates goes into `features`.
- The length check prevents a malformed row from reaching the classifier.
- Add the training workflow below `load_dataset()` by pasting this function:
def main():
args = parse_args()
features, labels = load_dataset(args.mode)
class_counts = Counter(labels)
if len(class_counts) != 3:
raise ValueError("Expected exactly three labels in the dataset.")
if min(class_counts.values()) < 4:
raise ValueError("Each label needs at least four samples.")
x_train, x_test, y_train, y_test = train_test_split(
features,
labels,
test_size=0.25,
random_state=42,
stratify=labels,
)
model = KNeighborsClassifier(n_neighbors=3, weights="uniform")
model.fit(x_train, y_train)
accuracy = model.score(x_test, y_test)
print(f"Holdout accuracy ({args.mode}): {accuracy:.1%}")
model.fit(features, labels)
MODEL_PATH.parent.mkdir(exist_ok=True)
with MODEL_PATH.open("wb") as model_file:
pickle.dump({"model": model, "mode": args.mode}, model_file)
print(f"Saved {args.mode} model to {MODEL_PATH}")
print(f"Class counts: {dict(sorted(class_counts.items()))}")
How Does Evaluation Become a Saved Model?
- `Counter` confirms that the dataset contains exactly three classes with enough samples.
- `train_test_split()` reserves 25% of the rows while preserving the class balance.
- `KNeighborsClassifier` uses three neighbors with equal voting weight.
- The second `fit()` call retrains the classifier on all 60 rows before saving the model and feature mode together.
- Finish `train_model.py` by adding this entry point at the bottom:
if __name__ == "__main__":
main()
Why Does the Entry Point Matter?
The entry point calls `main()` when you run `train_model.py` directly. This starts the load, evaluate, retrain, and save workflow.
✔️ Awesome, I've got everything!
Great. Double-check that you have saved `train_model.py`.
ⓧ I'd like to double check the full code
import argparse
from collections import Counter
import csv
from pathlib import Path
import pickle
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from hand_features import normalize_features
DATA_PATH = Path("data") / "landmarks.csv"
MODEL_PATH = Path("models") / "sign_knn.pkl"
def parse_args():
parser = argparse.ArgumentParser(description="Train the handshape classifier.")
parser.add_argument(
"--mode",
choices=("raw", "normalized"),
default="normalized",
)
return parser.parse_args()
def load_dataset(mode):
if not DATA_PATH.exists():
raise FileNotFoundError("Missing data/landmarks.csv. Collect samples first.")
features = []
labels = []
with DATA_PATH.open("r", newline="", encoding="utf-8") as data_file:
reader = csv.reader(data_file)
next(reader)
for row in reader:
sample = [float(value) for value in row[1:]]
if len(sample) != 63:
raise ValueError("Every dataset row must contain 63 feature values.")
if mode == "normalized":
sample = normalize_features(sample)
labels.append(row[0])
features.append(sample)
return features, labels
def main():
args = parse_args()
features, labels = load_dataset(args.mode)
class_counts = Counter(labels)
if len(class_counts) != 3:
raise ValueError("Expected exactly three labels in the dataset.")
if min(class_counts.values()) < 4:
raise ValueError("Each label needs at least four samples.")
x_train, x_test, y_train, y_test = train_test_split(
features,
labels,
test_size=0.25,
random_state=42,
stratify=labels,
)
model = KNeighborsClassifier(n_neighbors=3, weights="uniform")
model.fit(x_train, y_train)
accuracy = model.score(x_test, y_test)
print(f"Holdout accuracy ({args.mode}): {accuracy:.1%}")
model.fit(features, labels)
MODEL_PATH.parent.mkdir(exist_ok=True)
with MODEL_PATH.open("wb") as model_file:
pickle.dump({"model": model, "mode": args.mode}, model_file)
print(f"Saved {args.mode} model to {MODEL_PATH}")
print(f"Class counts: {dict(sorted(class_counts.items()))}")
if __name__ == "__main__":
main()
- Save `train_model.py`.
Before you train, which feature mode do you expect the saved payload to report?
- Train the raw-coordinate classifier by running this command in the terminal from earlier:
python train_model.py --mode raw
What Should the Training Report?
- The first line reports the raw-mode holdout accuracy.
- The next line confirms that the raw model was saved to `models\sign_knn.pkl`.
- The class counts confirm that `sign_1`, `sign_2`, and `sign_3` each contributed 20 samples.
Your exact accuracy depends on the examples you recorded. Treat it as a check on centered samples from the same recording conditions.
Training Did Not Create the Model?
- Confirm that `data\landmarks.csv` still contains its header plus 60 sample rows.
- Check that every row contains one label followed by 63 coordinate values.
- Confirm that your terminal still shows the activated `.venv` environment.
Still stuck? Help me diagnose why my raw handshape model did not train or save
Connect the saved model to the webcam
The saved payload lets a new process reuse the fitted classifier without recollecting data. The webcam loop now converts each detected hand into the same feature shape used during training.
- Create `live_predict.py` in the Visual Studio Code Explorer sidebar.
- Add the imports, model path, and payload loader by pasting this first section:
from pathlib import Path
import pickle
import cv2 as cv
from hand_features import (
create_landmarker,
detect_hand,
draw_landmarks,
normalize_features,
raw_features,
)
MODEL_PATH = Path("models") / "sign_knn.pkl"
def load_model():
if not MODEL_PATH.exists():
raise FileNotFoundError(
"Missing models/sign_knn.pkl. Run python train_model.py first."
)
with MODEL_PATH.open("rb") as model_file:
return pickle.load(model_file)
How Does the Predictor Restore the Model?
- `MODEL_PATH` points to the payload created by `train_model.py`.
- `load_model()` stops with a focused message when the saved classifier is missing.
- The loaded payload provides both the classifier and its feature mode.
- Add the webcam setup and prediction logic below `load_model()` by pasting this section:
def main():
payload = load_model()
model = payload["model"]
mode = payload["mode"]
capture = cv.VideoCapture(0)
if not capture.isOpened():
raise RuntimeError("Could not open webcam 0.")
try:
with create_landmarker() as landmarker:
while True:
frame_ok, frame = capture.read()
if not frame_ok:
raise RuntimeError("Could not read a frame from the webcam.")
landmarks = detect_hand(landmarker, frame)
if landmarks:
draw_landmarks(frame, landmarks)
features = raw_features(landmarks)
if mode == "normalized":
features = normalize_features(features)
prediction = str(model.predict([features])[0])
neighbor_vote = float(max(model.predict_proba([features])[0]))
result_text = f"{prediction} vote: {neighbor_vote:.0%}"
else:
result_text = "Show one hand"
How Does a Frame Become a Prediction?
- The webcam provides one frame at a time.
- `detect_hand()` returns one set of 21 landmarks.
- `raw_features()` converts those landmarks into 63 coordinate values.
- `predict_proba()` reveals the largest proportion of neighbors voting for one class.
- Continue `main()` by adding the display and webcam cleanup logic directly below the `else` section:
cv.putText(
frame,
result_text,
(20, 35),
cv.FONT_HERSHEY_SIMPLEX,
0.8,
(0, 255, 0),
2,
cv.LINE_AA,
)
cv.putText(
frame,
f"Feature mode: {mode}",
(20, 70),
cv.FONT_HERSHEY_SIMPLEX,
0.6,
(255, 255, 255),
2,
cv.LINE_AA,
)
cv.imshow("Personalized Sign Recognizer", frame)
if cv.waitKey(1) == ord("q"):
break
finally:
capture.release()
cv.destroyAllWindows()
What Does the Live Interface Show?
- The top line shows the predicted label beside the maximum neighbor vote.
- The second line shows the feature mode loaded from the saved payload.
- The `finally` block releases the webcam even when the loop stops because of an error.
- Finish `live_predict.py` by adding this entry point at the bottom:
if __name__ == "__main__":
main()
How Does the Live Loop Start?
The entry point calls `main()` when you run the file directly. This loads the saved payload before requesting the webcam.
✔️ Awesome, I've got everything!
Great. Double-check that you have saved `live_predict.py`.
ⓧ I'd like to double check the full code
from pathlib import Path
import pickle
import cv2 as cv
from hand_features import (
create_landmarker,
detect_hand,
draw_landmarks,
normalize_features,
raw_features,
)
MODEL_PATH = Path("models") / "sign_knn.pkl"
def load_model():
if not MODEL_PATH.exists():
raise FileNotFoundError(
"Missing models/sign_knn.pkl. Run python train_model.py first."
)
with MODEL_PATH.open("rb") as model_file:
return pickle.load(model_file)
def main():
payload = load_model()
model = payload["model"]
mode = payload["mode"]
capture = cv.VideoCapture(0)
if not capture.isOpened():
raise RuntimeError("Could not open webcam 0.")
try:
with create_landmarker() as landmarker:
while True:
frame_ok, frame = capture.read()
if not frame_ok:
raise RuntimeError("Could not read a frame from the webcam.")
landmarks = detect_hand(landmarker, frame)
if landmarks:
draw_landmarks(frame, landmarks)
features = raw_features(landmarks)
if mode == "normalized":
features = normalize_features(features)
prediction = str(model.predict([features])[0])
neighbor_vote = float(max(model.predict_proba([features])[0]))
result_text = f"{prediction} vote: {neighbor_vote:.0%}"
else:
result_text = "Show one hand"
cv.putText(
frame,
result_text,
(20, 35),
cv.FONT_HERSHEY_SIMPLEX,
0.8,
(0, 255, 0),
2,
cv.LINE_AA,
)
cv.putText(
frame,
f"Feature mode: {mode}",
(20, 70),
cv.FONT_HERSHEY_SIMPLEX,
0.6,
(255, 255, 255),
2,
cv.LINE_AA,
)
cv.imshow("Personalized Sign Recognizer", frame)
if cv.waitKey(1) == ord("q"):
break
finally:
capture.release()
cv.destroyAllWindows()
if __name__ == "__main__":
main()
- Save `live_predict.py`.
Before you run this, do you expect the label and neighbor vote to stay identical when one unchanged pose moves around the frame?
- Start the live raw-coordinate predictor by running this command in the terminal from earlier:
python live_predict.py
What Does This Command Start?
The command loads `models\sign_knn.pkl` before opening the webcam. The window shows a predicted label, its maximum neighbor vote, and `Feature mode: raw`.
- Hold your `sign_1` pose near the center of the frame.
- Repeat the centered test with `sign_2`.
- Repeat the centered test with `sign_3`.
- Keep one recognized pose unchanged.
- Move that pose toward the left edge of the frame.
- Move that pose toward the right edge of the frame.
- Move that pose closer to the webcam.
- Move that pose farther from the webcam.
You should see at least one predicted label or neighbor vote change during the movement test. That visible change is the intended limitation of absolute landmark coordinates.
You have now caught the naive classifier falling short in a live test. Its centered examples work best near the positions and sizes it already knows.
Cannot See the Raw-Coordinate Shortfall?
- Keep the pose itself steady while changing only its position in the image.
- Move close to the frame borders so the landmark coordinates change more strongly.
- Repeat the near and far test with each of your three poses.
Need help interpreting the result? Help me test whether my raw hand landmark classifier is position-sensitive
- Press `q` in the webcam window to release the camera.
Your classifier now predicts all three handshapes and exposes its sensitivity to position. Next, you will reshape the features so the same pose stays more stable as your hand moves.
Normalize the Hand and Stabilize Predictions
Your raw-coordinate k-nearest neighbors model recognized the centered poses. Its labels shifted when the same hand moved across the frame.
Wrist-relative translation removes the hand's image position. Scale normalization reduces the effect of hand size.
In this step, you will apply both transformations to every 63-value feature vector. You will retrain from the unchanged CSV before repeating the live movement test.
In this step, get ready to:
- Implement wrist-relative translation with scale normalization.
- Retrain the classifier with normalized features.
- Compare live predictions while moving the same pose.
Add hand-relative feature normalization
Each sample contains 21 points with three coordinates per point. The function uses the first point as the wrist reference.
The largest centered 3D distance becomes the scale for that sample. This produces a 63-value description based on hand shape.
- Select hand_features.py in the Visual Studio Code Explorer sidebar.
- Add this import above from pathlib import Path:
from math import sqrt
Why does normalization need this import?
The sqrt function calculates the 3D distance from the wrist to each centered landmark. The largest distance gives the hand its own scale.
- Scroll to the prepared normalize_features location.
- Replace the prepared implementation with this function:
def normalize_features(features):
if len(features) != 63:
raise ValueError("Expected 63 values from 21 three-dimensional landmarks.")
points = [features[index : index + 3] for index in range(0, 63, 3)]
wrist_x, wrist_y, wrist_z = points[0]
centered_points = [
(x - wrist_x, y - wrist_y, z - wrist_z) for x, y, z in points
]
scale = max(
sqrt(x * x + y * y + z * z) for x, y, z in centered_points
)
scale = max(scale, 1e-6)
return [coordinate / scale for point in centered_points for coordinate in point]
What does this function do?
- The length check protects the expected 63-value input shape.
- The point grouping rebuilds the 21 landmark positions.
- The wrist subtraction moves every hand to a shared origin.
- The maximum 3D distance measures the hand's scale in this sample.
- The 1e-6 minimum prevents division by zero.
- The final comprehension returns one flattened 63-value vector.
- Save hand_features.py.
✔️ Awesome, I've got everything!
Your helper now supports raw features plus normalized features.
ⓧ I'd like to double check the full code
Your complete hand_features.py file should match this reference.
from math import sqrt
from pathlib import Path
import time
import cv2 as cv
from mediapipe.tasks import python
from mediapipe.tasks.python import vision
from mediapipe.tasks.python.vision.core import image as image_lib
MODEL_PATH = Path("models") / "hand_landmarker.task"
_last_timestamp_ms = 0
def create_landmarker():
if not MODEL_PATH.exists():
raise FileNotFoundError(
"Missing models/hand_landmarker.task. Run python download_model.py first."
)
options = vision.HandLandmarkerOptions(
base_options=python.BaseOptions(model_asset_path=str(MODEL_PATH)),
running_mode=vision.RunningMode.VIDEO,
num_hands=1,
)
return vision.HandLandmarker.create_from_options(options)
def next_timestamp_ms():
global _last_timestamp_ms
current_timestamp_ms = time.monotonic_ns() // 1_000_000
_last_timestamp_ms = max(current_timestamp_ms, _last_timestamp_ms + 1)
return _last_timestamp_ms
def detect_hand(landmarker, frame):
rgb_frame = cv.cvtColor(frame, cv.COLOR_BGR2RGB)
image = image_lib.Image(
image_format=image_lib.ImageFormat.SRGB,
data=rgb_frame,
)
result = landmarker.detect_for_video(image, next_timestamp_ms())
if not result.hand_landmarks:
return None
return result.hand_landmarks[0]
def draw_landmarks(frame, landmarks):
height, width = frame.shape[:2]
for landmark in landmarks:
x = min(max(int(landmark.x * width), 0), width - 1)
y = min(max(int(landmark.y * height), 0), height - 1)
cv.circle(frame, (x, y), 4, (0, 255, 0), -1)
def raw_features(landmarks):
features = []
for landmark in landmarks:
features.extend(
(float(landmark.x), float(landmark.y), float(landmark.z))
)
return features
def normalize_features(features):
if len(features) != 63:
raise ValueError("Expected 63 values from 21 three-dimensional landmarks.")
points = [features[index : index + 3] for index in range(0, 63, 3)]
wrist_x, wrist_y, wrist_z = points[0]
centered_points = [
(x - wrist_x, y - wrist_y, z - wrist_z) for x, y, z in points
]
scale = max(
sqrt(x * x + y * y + z * z) for x, y, z in centered_points
)
scale = max(scale, 1e-6)
return [coordinate / scale for point in centered_points for coordinate in point]
This reference is the complete helper for normalized training. Live inference uses the same helper.
Retrain and compare live behavior
The existing train_model.py switches feature preparation based on its selected mode. Normalized mode transforms each row in memory.
The original observations remain unchanged in data\landmarks.csv. This keeps the raw experiment available for comparison.
- Retrain the classifier with normalized features by running:
python train_model.py --mode normalized
What does this training command change?
The script reloads all 60 raw rows. It calls normalize_features on each sample before creating the holdout split.
After evaluation, the script fits the classifier on all 60 normalized rows. The saved payload records normalized as its feature mode.
The terminal prints a normalized holdout accuracy. It also confirms that the normalized model was saved to models\sign_knn.pkl.
Training command not completing?
- Confirm the terminal prompt shows the active .venv environment.
- Check that from math import sqrt remains at the top of hand_features.py.
- Compare normalize_features with the full-file reference if the 63-value check interrupts training.
Still stuck? Help me debug normalized model training
The saved mode tells live_predict.py which feature path to apply. A normalized model receives normalized webcam input through the same function.
The webcam window keeps this terminal command running until you press q.
Before you start it, do you expect the same pose to keep a steadier label as its position changes?
- Start the normalized live recognizer by running:
python live_predict.py
What happens during normalized inference?
The application loads the saved classifier mode. It applies normalize_features before requesting a prediction.
The displayed vote remains the largest proportion returned by the three nearest neighbors. You can now compare its stability with the raw-model test.
- Hold each learned pose near the center of the frame.
- Move one unchanged pose toward the left edge.
- Move the same pose toward the right edge.
- Bring the same pose closer to the webcam.
- Move the same pose farther from the webcam.
- Repeat the movement sequence with sign_2.
- Repeat the movement sequence with sign_3.
- Press q to release the webcam.
You should see Feature mode: normalized in the webcam window. The predicted label should remain more stable across the movement test.
The neighbor vote should also fluctuate less than it did with the raw-coordinate model.
Predictions still changing too often?
- Return to the normalized training command above if the webcam window still shows raw mode.
- Keep the handshape fixed while testing one movement at a time.
- Use the same hand from your recorded dataset.
- Close any other application that is using the webcam.
Need another pair of eyes? Help me diagnose unstable normalized predictions
You have closed the gap exposed by the raw model. The same pose now holds its identity across more of the frame.
Secret mission
Reject Uncertain Predictions
Your recognizer currently chooses one of its three learned labels for every detected hand. Add a rejection rule that displays uncertain whenever the nearest neighbors disagree.
Clean Up Your Resources
Clean Up Your Resources
Everything in this project stays on your Windows computer, so there are no ongoing service fees. Choose whether to keep the local files, pause the active environment, or delete the project folder.
Resources you used:
- Local project folder open in Visual Studio Code with your Python source files and requirements.txt.
- Local .venv virtual environment containing the pinned MediaPipe, OpenCV, and scikit-learn packages.
- Model files models\hand_landmarker.task and models\sign_knn.pkl.
- Dataset data\landmarks.csv containing 60 labeled landmark rows.
Keep everything running
Keeping everything preserves your trained recognizer for future demonstrations. No service continues running after you close the webcam window.
- Press q in the webcam window if it is still open.
- Retain the current project folder with its environment, dataset, and models.
- Keep Visual Studio Code open if you plan to continue experimenting now.
Pause - I'll come back to this later
Pausing closes the webcam process and exits the active environment. Every project file remains available for your next session.
- Press q in the webcam window if it is still open.
- Deactivate the current environment by running this command:
deactivate
What does this command do?
The command exits the active virtual environment for this terminal session. It leaves the .venv folder and its installed packages on disk.
- Check the terminal prompt for the missing .venv prefix.
Your pause is complete. The dataset and models remain ready for your next session.
Delete - I don't want to use this again
Removing the project folder looks drastic because the deletion becomes permanent. It removes only this local project because you created no cloud resources or paid accounts.
- Press q in the webcam window if it is still open.
- Right-click the top-level project folder in Visual Studio Code's Explorer sidebar.
- Choose the context-menu option that reveals the folder in Windows File Explorer.
- Close Visual Studio Code using its window close control.
The file browser now gives you access to the project folder without an editor or webcam process holding its files open. Windows asks for confirmation before permanently removing it.
- Move up one level in File Explorer so the project folder itself is visible.
- Select the project folder.
- Press Shift+Delete to permanently remove the selected folder.
- Confirm the permanent deletion in the Windows prompt.
- Confirm that the project folder is no longer listed in the current File Explorer window.
Cleanup is complete. The virtual environment, model files, dataset, and source files are now removed.
Nice Work!
Nice Work!
You did it! Your personalized webcam recognizer now identifies three static handshapes from landmark geometry while rejecting non-unanimous predictions as uncertain.
You've learned how to:
- Build a real-time webcam pipeline with MediaPipe Hand Landmarker. Display 21 green landmark points through OpenCV while your hand moves.
- Collect a balanced labeled dataset with 20 examples per class. Use the neutral labels sign_1, sign_2, and sign_3. Turn each detected hand into a 63-value feature vector.
- Train a three-neighbor k-nearest neighbors classifier. Measure its performance with holdout evaluation. Expose the position sensitivity of raw coordinates. Improve live predictions with wrist-relative translation plus scale normalization.
- Secret Mission: Add a rejection rule that displays uncertain in red whenever the nearest-neighbor vote is not unanimous. Keep accepted unanimous predictions green.
Ready to quiz yourself?