Build a KNN Penguin Species Predictor

Build and test a KNN model that predicts penguin species from measurements.

Introduction

30 Second Summary

A penguin's bill, flippers, and body mass hold clues about its species. Those clues become harder to judge when each measurement uses a different scale.

In this project, you will build a Python KNN classifier that predicts one of three penguin species from four real measurements downloaded from Kaggle. You will judge its performance with five-fold cross-validation, held-out accuracy, and a saved confusion matrix.

What You'll Build

Picture running your finished program in Visual Studio Code while the terminal reveals how accurately four measurements identify a penguin's species.

By the end of this project, you'll have:

  • A repeatable data check that reports cleaned row counts, species totals, and each measurement's numeric range.
  • A fair model comparison that shows how feature scaling changes five-fold cross-validation accuracy.
  • A tuned species predictor with held-out accuracy, per-class precision and recall, and a saved confusion_matrix.png image you can inspect.
  • Secret Mission: An evidence-based error analysis that identifies the lowest-recall species, the most common confusion, and one focused next experiment.

Are there any prerequisites?

You need basic Python familiarity and a Kaggle account. Step 1 guides you through installing or verifying Python, Visual Studio Code, and the VS Code Python extension.

Before We Start

Before any hands-on work, you'll define the practical question behind this project. This commitment gives every later modeling decision a clear purpose.

Set Up the Windows Project

The penguin classifier needs a reproducible local environment before any measurements reach KNN. Isolation keeps this experiment from borrowing package versions from another project.

A successful starter run proves that Visual Studio Code can use the selected Python environment. You will finish this step with the project structure ready for the real dataset inspection.

In this step, get ready to:
  • Prepare the required Windows development tools.
  • Create an isolated project environment with pinned dependencies.
  • Prove the project can run its starter script.
Verify the Windows tools

Python runs the experiment on your computer. Checking the installed version first tells you whether the required interpreter is ready.

  • Press the Windows key to open Windows search.
  • Type Terminal into the search bar.
  • Press Enter to open a Windows terminal.
  • Check the installed Python 3 version by running this command:
py -3 --version

What does this command do?

The command asks the Windows Python launcher to report its Python 3 version. The printed version determines which setup path you need.

✔️ I see version 3.14.8 or higher

Your Python version meets this project's requirement. Keep the terminal available for the dependency installation later in this step.

ⓧ I see an older version

Your current Python version is below the project's tested version. Install Python 3.14.8 before creating the environment.

You are ready when the terminal reports Python 3.14.8 or higher.

ⓧ Command not found

Windows cannot find a Python installation yet. Install the required version from Python's official page.

You are ready when the terminal prints an installed Python 3 version.

Visual Studio Code provides the workspace for your files. Its Python extension creates the project environment and runs the selected script.

Create the isolated project

A virtual environment keeps this project's packages separate from other Python work. The version pins make the experiment repeatable on the tested package releases.

  • Press the Windows key to open Windows search.
  • Type Visual Studio Code into the search bar.
  • Press Enter to open Visual Studio Code.
  • Click File in the top menu.
  • Select Open Folder.
  • Navigate to your Desktop in the folder picker.
  • Create a folder named knn-penguins on your Desktop.
  • Select knn-penguins to open it as your Visual Studio Code workspace.

You should see knn-penguins in the Explorer sidebar. That folder is now the home for every local project artifact.

The requirements file records the exact library versions used by the experiment. Installing from this file gives every run the same dependency targets.

  • Create requirements.txt inside knn-penguins from the Explorer sidebar.
  • Pin the project dependencies by copying this code into requirements.txt:
pandas==3.0.6
scikit-learn==1.9.1
matplotlib==3.11.2

What are these packages?

  • pandas 3.0.6 loads the penguin measurements from the CSV file.
  • scikit-learn 1.9.1 supplies the KNN model. Later steps use its evaluation tools.
  • Matplotlib 3.11.2 saves the final confusion matrix as an image.
  • Save requirements.txt.
  • Open the Visual Studio Code Command Palette.
  • Enter Python: Create Environment.

You will see the environment types supported by the Python extension.

  • Select Venv.

You will see the available Python interpreters.

  • Select the Python 3.14.8 interpreter.

VS Code creates the Venv environment and selects it for this workspace. This can take a moment while the isolated environment is prepared.

  • Open the Visual Studio Code Command Palette again.
  • Enter Python: Select Interpreter.
  • Confirm that the new Venv environment is selected.

The selected environment confirms that Visual Studio Code runs this project with its isolated interpreter.

  • Open a new integrated terminal in Visual Studio Code.
  • Install the pinned dependencies from requirements.txt by running this command:
pip install -r requirements.txt

What does this command do?

The command reads each version pin from requirements.txt. It installs those packages into the selected Venv environment.

When installation finishes, the terminal returns to a prompt without an installation error. Your isolated environment now contains all three project libraries.

Having trouble installing the packages?

  • Confirm that Visual Studio Code still shows the Venv environment as the selected Python interpreter.
  • Check that requirements.txt contains the three version pins exactly as shown above.
  • Confirm that your computer has an active internet connection before trying the installation again.

Still stuck? Help me troubleshoot my pinned Python dependency installation.

Finish the runnable project

The real measurements come from the Kaggle Palmer Penguins dataset. The downloaded CSV must live at data/penguins.csv so the next step can load it from a stable path.

The Kaggle download can be a little fiddly because it may arrive inside an archive. Extract the file before placing it in your project.

  • Open the Palmer Penguins dataset page.
  • Download the Palmer Penguins dataset to your Windows Downloads folder.
  • Extract penguins.csv if the download arrives as an archive.
  • Create a folder named data inside knn-penguins from the Visual Studio Code Explorer sidebar.
  • Drag penguins.csv from your Downloads folder into the data folder in the Explorer sidebar.

You should now see the downloaded file at data/penguins.csv in the Explorer sidebar. That path confirms the dataset is ready for the next step.

  • Create knn_penguins.py inside knn-penguins from the Explorer sidebar.
  • Add the runnable checkpoint by copying this code into knn_penguins.py:
print("KNN project ready")

What does this code do?

The print statement gives you a visible checkpoint before the data logic is added. A successful message proves that Visual Studio Code can run the selected Python file.

  • Save knn_penguins.py.

Before you run the script, what message do you expect the selected Python environment to print?

  • Select Run Python File in the editor toolbar.

You should see KNN project ready in the Visual Studio Code terminal. That is your first working checkpoint for the local project.

Don't see the ready message?

  • Confirm that knn_penguins.py is the open file in the editor.
  • Save knn_penguins.py before selecting Run Python File again.
  • Confirm that the VS Code Python extension is installed if the run control is missing.

Need another pair of eyes? Help me run my Python starter file in Visual Studio Code.

✔️ Awesome, I've got everything!

Your dataset path and project files are in place. The successful terminal message proves that the selected environment can run the starter script.

ⓧ I'd like to double check the full code

  • Confirm that the downloaded Kaggle Palmer Penguins CSV appears at data/penguins.csv.
  • Compare requirements.txt with this complete reference:
pandas==3.0.6
scikit-learn==1.9.1
matplotlib==3.11.2

What should this file contain?

This file contains the exact pandas, scikit-learn, and Matplotlib version pins for the project. Each dependency appears on its own line.

  • Compare knn_penguins.py with this complete reference:
print("KNN project ready")

What should this file do?

This starter file prints the project readiness message. The next step replaces this checkpoint with the data inspection code.

Your Windows project is ready to run with its pinned dependencies. Next up, you will load the penguin measurements and expose their mismatched numeric ranges.

Inspect the Penguin Data

Your starter script already proves that the selected Python environment in Visual Studio Code can run your project. The next question is whether the downloaded rows are ready for distance-based learning.

K-nearest neighbors classification compares numeric distances between penguins. Features with larger numeric ranges can dominate those distances.

In this step, get ready to:
  • Define the dataset inputs used by the experiment.
  • Build a complete modeling table from the downloaded CSV.
  • Expose the class balance plus feature ranges in the terminal.
Load and clean the modeling table

pandas reads penguins.csv into a DataFrame. Keeping the required columns gives every modeling row the same four measurements plus its species label.

  • Return to knn_penguins.py in the VS Code editor.
  • Select the starter line print("KNN project ready").
  • Replace the selected line by pasting this code:
import pandas as pd
from sklearn.metrics import (
    ConfusionMatrixDisplay,
    accuracy_score,
    classification_report,
)
from sklearn.model_selection import GridSearchCV, cross_val_score, train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler

DATA_PATH = "data/penguins.csv"
FEATURES = [
    "bill_length_mm",
    "bill_depth_mm",
    "flipper_length_mm",
    "body_mass_g",
]
TARGET = "species"

# Load only the columns needed for this classification experiment.
df = pd.read_csv(DATA_PATH)
model_data = df[FEATURES + [TARGET]].dropna()

print(f"Downloaded rows: {len(df)}")
print(f"Rows used after dropping missing values: {len(model_data)}")
print("\nClass counts:")
print(model_data.value_counts(subset=TARGET))

What does this code do?

  • The pandas import provides pd.read_csv() for loading the downloaded CSV into df.
  • The scikit-learn imports prepare the modeling tools used in later steps.
  • DATA_PATH points to the downloaded file inside the data folder.
  • FEATURES identifies the four physical measurements used to compare penguins.
  • TARGET identifies the species label that the finished model predicts.
  • dropna() removes rows missing any required modeling value.
  • value_counts() reports how many complete examples belong to each species.
  • Save knn_penguins.py.

Before you run the script, do you expect every downloaded row to remain in the cleaned modeling table?

  • Run the inspection by clicking Run Python File in the top-right corner of the editor.

You should see the downloaded row count followed by the smaller cleaned row count. The class-count section should list Adelie, Gentoo, and Chinstrap.

Can't load the CSV?

  • Check that the data folder sits directly inside knn-penguins.
  • Confirm that the downloaded file is named penguins.csv.
  • Extract the downloaded archive if Kaggle supplied the CSV inside one.

Still stuck? Help me find why my penguin CSV is not loading.

Interpret the row and class checks

Why remove incomplete rows?

This experiment represents each penguin with the same four measurements. Dropping incomplete rows keeps every example aligned to that input shape.

The two row counts make the cleanup visible. The species counts reveal how the usable examples are distributed across the three classes.

  • Compare the terminal lines beginning with Downloaded rows: and Rows used after dropping missing values:.
  • Confirm that the Class counts: section contains three species.

A lower cleaned count confirms that incomplete examples were excluded. Three class counts confirm that every target species remains represented.

Reveal the feature ranges

A feature range is the difference between its maximum value and minimum value. Comparing those ranges reveals how strongly the gram measurements outweigh the millimeter measurements numerically.

  • In knn_penguins.py, find print(model_data.value_counts(subset=TARGET)).
  • Add the range diagnostic directly below that line by pasting this code:
print("\nFeature ranges before scaling:")
for feature in FEATURES:
    minimum = min(model_data[feature])
    maximum = max(model_data[feature])
    print(
        f"{feature}: min={minimum:.1f}, "
        f"max={maximum:.1f}, range={maximum - minimum:.1f}"
    )

What does this code do?

  • The loop inspects every measurement named in FEATURES.
  • minimum stores the smallest observed value for the current feature.
  • maximum stores the largest observed value for the current feature.
  • The formatted output displays each minimum to one decimal place.
  • The same output displays each maximum to one decimal place.
  • The final calculation exposes the numeric range used by distance comparisons.
  • Save knn_penguins.py.

Before you run the finished inspection, which measurement do you expect to have the widest numeric range?

  • Run the finished inspection by clicking Run Python File.

The terminal should print the downloaded row count first. It should then print the cleaned row count, three species counts, and four feature ranges.

You should see a much wider numeric range for body_mass_g than for the millimeter measurements. That visible mismatch proves raw distance calculations can give the gram-based feature excessive influence.

Don't see the feature ranges?

  • Check that the new code sits below print(model_data.value_counts(subset=TARGET)).
  • Match the indentation beneath for feature in FEATURES: to the completed code.
  • Confirm that every feature name in FEATURES matches the CSV column spelling.

Need another pair of eyes? Help me debug my penguin feature-range output.

✔️ Awesome, I've got everything!

Great work. Your script now loads the Kaggle data, removes incomplete modeling rows, and exposes the feature-scale mismatch.

ⓧ I'd like to double check the full code

  • Compare your saved knn_penguins.py file with this complete version:
import pandas as pd
from sklearn.metrics import (
    ConfusionMatrixDisplay,
    accuracy_score,
    classification_report,
)
from sklearn.model_selection import GridSearchCV, cross_val_score, train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler

DATA_PATH = "data/penguins.csv"
FEATURES = [
    "bill_length_mm",
    "bill_depth_mm",
    "flipper_length_mm",
    "body_mass_g",
]
TARGET = "species"

# Load only the columns needed for this classification experiment.
df = pd.read_csv(DATA_PATH)
model_data = df[FEATURES + [TARGET]].dropna()

print(f"Downloaded rows: {len(df)}")
print(f"Rows used after dropping missing values: {len(model_data)}")
print("\nClass counts:")
print(model_data.value_counts(subset=TARGET))

print("\nFeature ranges before scaling:")
for feature in FEATURES:
    minimum = min(model_data[feature])
    maximum = max(model_data[feature])
    print(
        f"{feature}: min={minimum:.1f}, "
        f"max={maximum:.1f}, range={maximum - minimum:.1f}"
    )

Your clean modeling table now makes the distance problem visible. Next up, you'll measure how a naive KNN classifier performs before correcting that mismatch.

Measure an Unscaled KNN Baseline

Your feature-range check showed that body mass spans much larger numbers than the measurements recorded in millimeters. Now you need a baseline that captures how the model performs with those raw values.

A KNN classifier predicts a species using the labels of nearby training examples. This deliberately naive experiment measures closeness before correcting the unequal feature ranges.

In this step, get ready to:
  • Separate the cleaned rows into training and held-out test partitions.
  • Train a five-neighbor classifier on the raw feature values.
  • Measure accuracy with five-fold cross-validation using only the training partition.
Create the training and test partitions

A train-test split reserves part of the cleaned data for a final evaluation. The training partition handles every decision made before that final check.

  • In the open knn_penguins.py file, scroll to the feature-range loop at the bottom.
  • Find the closing parenthesis beneath the formatted feature-range output.
  • Paste this code on a new line below that parenthesis:
X = model_data[FEATURES]
y = model_data[TARGET]

X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.20,
    random_state=42,
    stratify=y,
)

What does this code do?

  • X stores the four physical measurements used to make predictions.
  • y stores the species label that the model learns to predict.
  • test_size=0.20 reserves 20 percent of the cleaned rows for the held-out test partition.
  • random_state=42 makes the split reproducible across repeated runs.
  • stratify=y preserves the species mix across the training and test partitions.
  • Save knn_penguins.py.
  • Select Run Python File in the top-right corner of Visual Studio Code.

You should still see the downloaded row count, cleaned row count, species counts, and feature ranges. That output confirms the new split runs without interrupting the existing data checks.

Does the script stop before printing the ranges?

  • Check that X and y appear after the complete feature-range loop.
  • Confirm that every new line begins at the far-left edge of knn_penguins.py except the arguments inside train_test_split().

Still stuck? Help me check why my train-test split stops the penguin script.

Measure the five-neighbor baseline

Five-fold cross-validation divides the training partition into five folds. Each round trains on four folds before measuring accuracy on the remaining fold.

  • Return to the bottom of knn_penguins.py.
  • Paste this code beneath the closing parenthesis of train_test_split():
# Establish a deliberately naive baseline before fixing feature scale.
baseline_model = KNeighborsClassifier(n_neighbors=5)
baseline_scores = cross_val_score(
    baseline_model,
    X_train,
    y_train,
    cv=5,
)
baseline_accuracy = sum(baseline_scores) / len(baseline_scores)
print(f"\nUnscaled 5-fold CV accuracy: {baseline_accuracy:.3f}")

What does this code do?

  • baseline_model creates a classifier that considers the five nearest training examples.
  • cross_val_score() evaluates that classifier across five folds of the training partition.
  • X_test and y_test stay outside the cross-validation process.
  • baseline_accuracy averages the five fold scores into one comparison value.
  • .3f formats the average with three digits after the decimal point.
  • Save knn_penguins.py.

Before you run the script, do you think the raw feature ranges give every measurement equal influence over which penguins count as neighbors?

  • Select Run Python File in the top-right corner of Visual Studio Code.

You'll see Unscaled 5-fold CV accuracy: followed by a numeric score between 0 and 1. That score is your repeatable reference point.

The model completed the experiment on raw units. Its neighbor distances give larger-number measurements more influence.

This is the intended shortfall exposed by the earlier feature-range check. The held-out test partition remains untouched for the final evaluation.

Don't see the baseline accuracy?

  • Check that baseline_scores receives X_train and y_train.
  • Confirm that cv=5 sits inside the cross_val_score() call.
  • Compare the spelling of baseline_accuracy in the calculation with its spelling in the final print line.

Need another pair of eyes? Help me troubleshoot my missing unscaled cross-validation score.

Your cumulative knn_penguins.py file should now match one of the outcomes below.

✔️ Awesome, I've got everything!

Baseline complete. Your saved script now creates a reproducible split and measures five-fold accuracy without using the held-out test partition.

ⓧ I'd like to double check the full code

import pandas as pd
from sklearn.metrics import (
    ConfusionMatrixDisplay,
    accuracy_score,
    classification_report,
)
from sklearn.model_selection import GridSearchCV, cross_val_score, train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler

DATA_PATH = "data/penguins.csv"
FEATURES = [
    "bill_length_mm",
    "bill_depth_mm",
    "flipper_length_mm",
    "body_mass_g",
]
TARGET = "species"

# Load only the columns needed for this classification experiment.
df = pd.read_csv(DATA_PATH)
model_data = df[FEATURES + [TARGET]].dropna()

print(f"Downloaded rows: {len(df)}")
print(f"Rows used after dropping missing values: {len(model_data)}")
print("\nClass counts:")
print(model_data.value_counts(subset=TARGET))

print("\nFeature ranges before scaling:")
for feature in FEATURES:
    minimum = min(model_data[feature])
    maximum = max(model_data[feature])
    print(
        f"{feature}: min={minimum:.1f}, "
        f"max={maximum:.1f}, range={maximum - minimum:.1f}"
    )

X = model_data[FEATURES]
y = model_data[TARGET]

X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.20,
    random_state=42,
    stratify=y,
)

# Establish a deliberately naive baseline before fixing feature scale.
baseline_model = KNeighborsClassifier(n_neighbors=5)
baseline_scores = cross_val_score(
    baseline_model,
    X_train,
    y_train,
    cv=5,
)
baseline_accuracy = sum(baseline_scores) / len(baseline_scores)
print(f"\nUnscaled 5-fold CV accuracy: {baseline_accuracy:.3f}")

You now have a repeatable baseline built only from the training partition. Next, you'll standardize the feature values and compare both workflows on the same training data.

Scale the KNN Features

Your unscaled KNN cross-validation score now provides a baseline. Its distance calculation still lets body mass dominate because grams span a much larger numeric range than the millimeter measurements.

Scaling puts every feature on comparable footing before the model measures neighbors. In this step, you will place StandardScaler with KNeighborsClassifier inside one Pipeline before comparing the same five-fold process.

In this step, get ready to:
  • Build a scaled pipeline that standardizes every measurement before KNN calculates distance.
  • Cross-validate the complete workflow using only the training partition.
  • Compare the scaled accuracy with the unscaled baseline.
Scale Features

A scaler must learn from each training fold before transforming that fold's validation rows. Keeping both operations in one pipeline preserves that order during every evaluation.

  • Find the final accuracy print line in knn_penguins.py.
  • Add the scaled workflow directly below that line by copying this code:
# Keep scaling inside the pipeline so every fit uses training-fold statistics.
scaled_model = Pipeline(
    steps=[
        ("scaler", StandardScaler()),
        ("knn", KNeighborsClassifier(n_neighbors=5)),
    ]
)
scaled_scores = cross_val_score(
    scaled_model,
    X_train,
    y_train,
    cv=5,
)
scaled_accuracy = sum(scaled_scores) / len(scaled_scores)
print(f"Scaled 5-fold CV accuracy: {scaled_accuracy:.3f}")

Why Keep Scaling in the Pipeline?

  • The pipeline stores preprocessing with the classifier as one workflow.
  • The scaler step uses StandardScaler() to standardize the four measurements.
  • The knn step keeps the same five-neighbor classifier used by the baseline.
  • The cross_val_score() call evaluates the complete scaled workflow across five folds of the training partition.
  • The mean of scaled_scores produces one accuracy value that you can compare with baseline_accuracy.
  • Save knn_penguins.py.

✔️ Awesome, I've got everything!

  • Confirm that knn_penguins.py is saved with the scaled workflow below the baseline.

ⓧ I'd like to double check the full code

  • Compare your saved knn_penguins.py with this cumulative version:
import pandas as pd
from sklearn.metrics import (
    ConfusionMatrixDisplay,
    accuracy_score,
    classification_report,
)
from sklearn.model_selection import GridSearchCV, cross_val_score, train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler

DATA_PATH = "data/penguins.csv"
FEATURES = [
    "bill_length_mm",
    "bill_depth_mm",
    "flipper_length_mm",
    "body_mass_g",
]
TARGET = "species"

# Load only the columns needed for this classification experiment.
df = pd.read_csv(DATA_PATH)
model_data = df[FEATURES + [TARGET]].dropna()

print(f"Downloaded rows: {len(df)}")
print(f"Rows used after dropping missing values: {len(model_data)}")
print("\nClass counts:")
print(model_data.value_counts(subset=TARGET))

print("\nFeature ranges before scaling:")
for feature in FEATURES:
    minimum = min(model_data[feature])
    maximum = max(model_data[feature])
    print(
        f"{feature}: min={minimum:.1f}, "
        f"max={maximum:.1f}, range={maximum - minimum:.1f}"
    )

X = model_data[FEATURES]
y = model_data[TARGET]

X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.20,
    random_state=42,
    stratify=y,
)

# Establish a deliberately naive baseline before fixing feature scale.
baseline_model = KNeighborsClassifier(n_neighbors=5)
baseline_scores = cross_val_score(
    baseline_model,
    X_train,
    y_train,
    cv=5,
)
baseline_accuracy = sum(baseline_scores) / len(baseline_scores)
print(f"\nUnscaled 5-fold CV accuracy: {baseline_accuracy:.3f}")

# Keep scaling inside the pipeline so every fit uses training-fold statistics.
scaled_model = Pipeline(
    steps=[
        ("scaler", StandardScaler()),
        ("knn", KNeighborsClassifier(n_neighbors=5)),
    ]
)
scaled_scores = cross_val_score(
    scaled_model,
    X_train,
    y_train,
    cv=5,
)
scaled_accuracy = sum(scaled_scores) / len(scaled_scores)
print(f"Scaled 5-fold CV accuracy: {scaled_accuracy:.3f}")

What Should Match?

This reference contains the complete script for the project so far. The final section starts with the scaling comment before defining scaled_model.

Compare Results

Both workflows use only X_train with y_train. Their five-fold averages create a direct comparison without evaluating the held-out test partition.

Before you run the script, do you expect scaling to change the cross-validation score?

  • Select Run Python File from the top-right editor toolbar.

Your terminal prints both Unscaled 5-fold CV accuracy: and Scaled 5-fold CV accuracy: with numeric scores. You have now made the model compare penguin measurements on compatible scales.

Missing the Scaled Accuracy?

  • Confirm that the scaled workflow appears below the line that prints Unscaled 5-fold CV accuracy:.
  • Save knn_penguins.py before selecting Run Python File again.
  • Compare the parentheses around Pipeline() with the full-code reference if Python reports a syntax problem.

Still stuck? Help me debug why my scaled KNN accuracy is missing.

Why Scaling Still Matters

Small datasets can produce similar cross-validation averages. Scaling remains the correct KNN workflow because each standardized feature contributes on a comparable numeric scale.

This protects neighbor selection from body_mass_g dominating solely through its larger numbers.

Your scaled workflow now gives KNN a fair distance calculation while preserving the unscaled baseline for comparison. Next, you will tune the number of neighbors before evaluating the untouched test data.

Tune K and Test the Model

Your scaled KNN workflow now compares every measurement on a consistent scale. The next goal is to choose the neighbor count that performs best during cross-validation.

A fixed value of k is only a starting guess. Grid search compares several candidate values without using the held-out test set.

The held-out test set has stayed untouched so far. It now measures how the selected model performs on unseen penguins.

In this step, get ready to:
  • Tune the scaled pipeline across odd values of k from 1 through 21.
  • Evaluate the selected model on held-out data with accuracy plus a classification report.
  • Save a confusion matrix that shows the model's correct predictions plus its mistakes.
Tune k and evaluate held-out data

The GridSearchCV process evaluates each candidate through five folds of the training partition. Its selected pipeline can then predict the untouched test partition.

  • Find the final print(f"Scaled 5-fold CV accuracy: {scaled_accuracy:.3f}") line in knn_penguins.py:
  • Add the tuning plus evaluation logic below that line by copying this code:
# Tune k by cross-validation using only the training partition.
parameter_grid = {
    "knn__n_neighbors": list(range(1, 22, 2)),
}
search = GridSearchCV(
    estimator=scaled_model,
    param_grid=parameter_grid,
    cv=5,
)
search.fit(X_train, y_train)

final_predictions = search.predict(X_test)
final_accuracy = accuracy_score(y_test, final_predictions)

print(f"\nSelected parameters: {search.best_params_}")
print(f"Best 5-fold CV accuracy: {search.best_score_:.3f}")
print(f"Held-out test accuracy: {final_accuracy:.3f}")
print("\nClassification report:")
print(classification_report(y_test, final_predictions))

What does this code do?

  • The parameter_grid stores every odd value of k from 1 through 21. The knn__n_neighbors name targets the neighbor setting inside the pipeline's named knn step.
  • The search object evaluates those candidates with five-fold cross-validation. search.fit(X_train, y_train) keeps that selection process inside the training partition.
  • The final_predictions variable stores predictions for the held-out penguins. The final_accuracy variable stores the fraction predicted correctly.
  • The final print statements expose the selected parameters. They also report cross-validation accuracy plus held-out accuracy.
  • The classification_report output breaks the test results into precision plus recall for each species.
  • Save knn_penguins.py.

Before you run the script, which odd value of k do you think the training folds will select?

  • Select Run Python File in the editor to run the updated script.

The terminal now prints Selected parameters: with the chosen neighbor setting. It also prints Best 5-fold CV accuracy: with the training-only result.

You will then see Held-out test accuracy: with a score between 0 and 1. A classification report follows for the three species.

You have now tuned the model without using the test partition for model selection. The final accuracy therefore measures performance on unseen examples.

Missing the final metrics?

  • Check that the new code sits below the scaled cross-validation output at the leftmost indentation level.
  • Confirm that knn__n_neighbors contains two underscore characters between the pipeline step name plus the parameter name.
  • Compare the parentheses around GridSearchCV with the code block above if the script stops before fitting the search.

Still stuck? Help me debug the grid search and held-out evaluation in my penguin KNN script.

Save and inspect the confusion matrix

A confusion matrix arranges the actual species against the model's predicted species. Its off-diagonal cells reveal which species the model confused.

  • Find the final print(classification_report(y_test, final_predictions)) line in knn_penguins.py:
  • Add the matrix creation plus image-saving logic below that line by copying this code:
confusion_display = ConfusionMatrixDisplay.from_predictions(
    y_test,
    final_predictions,
)
confusion_display.figure_.savefig("confusion_matrix.png")
print("\nSaved confusion_matrix.png")

How is the matrix saved?

  • The from_predictions method compares the true test labels with the model's predictions. It turns those results into a labeled matrix.
  • The figure_ attribute provides access to the Matplotlib figure containing the matrix.
  • The savefig call writes that figure to confusion_matrix.png inside the knn-penguins folder.
  • The final print statement confirms that the image-saving line completed.
  • Save knn_penguins.py.

Before you run the finished script, where do you expect the largest counts to appear if most predictions are correct?

  • Select Run Python File in the editor to run the finished script.

The terminal prints the selected k plus its cross-validation accuracy. It then prints Held-out test accuracy: plus the classification report.

The final terminal line reads Saved confusion_matrix.png. You will also see confusion_matrix.png in Visual Studio Code's project file list.

  • Select confusion_matrix.png from the project file list to inspect the model's predictions.

You should see a labeled grid with counts for each actual species plus each predicted species. Cells away from the main diagonal expose the model's mistakes.

That completes your KNN experiment. You now have training-only model selection plus an independent test evaluation you can inspect visually.

Missing the matrix image?

  • Check that the matrix code sits below the classification report at the leftmost indentation level.
  • Confirm that the filename is exactly confusion_matrix.png if the terminal confirms a save but the expected file is missing.
  • Run the complete script again if the project file list still shows results from the earlier version.

Need another pair of eyes? Help me find why my KNN script is not saving or displaying confusion_matrix.png.

✔️ Awesome, I've got everything!

Your saved script now tunes k on the training partition. It reports the held-out results plus saves confusion_matrix.png.

ⓧ I'd like to double check the full code

  • Compare your knn_penguins.py file against this complete version:
import pandas as pd
from sklearn.metrics import (
    ConfusionMatrixDisplay,
    accuracy_score,
    classification_report,
)
from sklearn.model_selection import GridSearchCV, cross_val_score, train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler

DATA_PATH = "data/penguins.csv"
FEATURES = [
    "bill_length_mm",
    "bill_depth_mm",
    "flipper_length_mm",
    "body_mass_g",
]
TARGET = "species"

# Load only the columns needed for this classification experiment.
df = pd.read_csv(DATA_PATH)
model_data = df[FEATURES + [TARGET]].dropna()

print(f"Downloaded rows: {len(df)}")
print(f"Rows used after dropping missing values: {len(model_data)}")
print("\nClass counts:")
print(model_data.value_counts(subset=TARGET))

print("\nFeature ranges before scaling:")
for feature in FEATURES:
    minimum = min(model_data[feature])
    maximum = max(model_data[feature])
    print(
        f"{feature}: min={minimum:.1f}, "
        f"max={maximum:.1f}, range={maximum - minimum:.1f}"
    )

X = model_data[FEATURES]
y = model_data[TARGET]

X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.20,
    random_state=42,
    stratify=y,
)

# Establish a deliberately naive baseline before fixing feature scale.
baseline_model = KNeighborsClassifier(n_neighbors=5)
baseline_scores = cross_val_score(
    baseline_model,
    X_train,
    y_train,
    cv=5,
)
baseline_accuracy = sum(baseline_scores) / len(baseline_scores)
print(f"\nUnscaled 5-fold CV accuracy: {baseline_accuracy:.3f}")

# Keep scaling inside the pipeline so every fit uses training-fold statistics.
scaled_model = Pipeline(
    steps=[
        ("scaler", StandardScaler()),
        ("knn", KNeighborsClassifier(n_neighbors=5)),
    ]
)
scaled_scores = cross_val_score(
    scaled_model,
    X_train,
    y_train,
    cv=5,
)
scaled_accuracy = sum(scaled_scores) / len(scaled_scores)
print(f"Scaled 5-fold CV accuracy: {scaled_accuracy:.3f}")

# Tune k by cross-validation using only the training partition.
parameter_grid = {
    "knn__n_neighbors": list(range(1, 22, 2)),
}
search = GridSearchCV(
    estimator=scaled_model,
    param_grid=parameter_grid,
    cv=5,
)
search.fit(X_train, y_train)

final_predictions = search.predict(X_test)
final_accuracy = accuracy_score(y_test, final_predictions)

print(f"\nSelected parameters: {search.best_params_}")
print(f"Best 5-fold CV accuracy: {search.best_score_:.3f}")
print(f"Held-out test accuracy: {final_accuracy:.3f}")
print("\nClassification report:")
print(classification_report(y_test, final_predictions))

confusion_display = ConfusionMatrixDisplay.from_predictions(
    y_test,
    final_predictions,
)
confusion_display.figure_.savefig("confusion_matrix.png")
print("\nSaved confusion_matrix.png")

Secret mission

Investigate the Model's Mistakes

Your final accuracy score summarizes the test set, but it does not reveal which species is hardest to identify. Use the classification report and confusion matrix to write a short evidence-based model review.

Clean Up Your Resources

Clean Up Your Resources

Your KNN experiment runs entirely on your Windows computer. It has no ongoing cloud costs, so you can keep your resources available, pause your work, or delete them entirely.

Resources you used:

  • Local Visual Studio Code project folder knn-penguins, including knn_penguins.py, requirements.txt, data/penguins.csv, confusion_matrix.png, model_review.md, and the selected Venv.

Keep everything running

No action needed. Choose this if you want to rerun the experiment or continue testing model improvements.

  • Keep the knn-penguins folder in its current location.
  • Keep the selected Venv with the project so the pinned packages stay available.
  • Return to knn_penguins.py whenever you want to rerun the experiment.

Pause - I'll come back to this later

Pausing means closing the open workspace to free up memory while keeping every project file. The script has no background server to stop.

  • Save knn_penguins.py in VS Code.
  • Save model_review.md in VS Code.
  • Close the open VS Code window.
  • Leave the knn-penguins folder in its current location.
  • Reopen the existing knn-penguins folder in VS Code when you return.

Delete - I don't want to use this again

Deleting knn-penguins removes the full local experiment in one step. This only affects the selected project folder on your computer.

  • Copy model_review.md to another folder if you want to keep your written findings.
  • Copy confusion_matrix.png to another folder if you want to keep the visual evaluation.
  • Close the open VS Code window.
  • Press the Windows key to open search.
  • Search for File Explorer.
  • Select File Explorer from the search results.
  • Browse to the location that contains your knn-penguins folder.
  • Select the knn-penguins folder.
  • Press Shift+Delete to permanently delete the selected folder.
  • Approve the confirmation prompt if Windows displays one.

You should no longer see knn-penguins in its previous location. The script, dataset, generated image, written review, pinned requirements, and project Venv are now removed.

Nice Work!

Nice Work!

You did it! Your local Python experiment now uses KNN classification to predict penguin species from four physical measurements.

You've learned how to:

  • Prepared the real Kaggle dataset with pandas so every modeling row has all four physical measurements plus its species label.
  • Proved why feature scaling belongs in distance-based learning by comparing unscaled and scaled KNN workflows with five-fold cross-validation.
  • Tuned k with GridSearchCV using only the training partition. Measured held-out accuracy on unseen data. Saved a confusion matrix as confusion_matrix.png for visual inspection.
  • Completed the optional Secret Mission by turning the evaluation results into model_review.md. The review names the lowest-recall species. It identifies the largest off-diagonal confusion. It proposes one focused next experiment.

Ready to quiz yourself?